Version 1 · written 2026-09-10 against plugrl-protocol@22b324c,
plugrl-server@06e9f63 and plugrl-env-client@b9edc5b.
This document specifies the bytes exchanged between a training server, which holds the policy and the learning algorithm, and an env client, which runs environments and asks for actions. Anything able to speak WebSocket and msgpack can be an env client: a Python process, a ROS node, a robot's onboard C++ controller.
The specification is written from the implementation, not the other way round. Where the implementation is under-specified, ambiguous, or where two of its own servers disagree, this document says so in a box marked Gap rather than inventing a rule. Those boxes are the honest part; treat them as the protocol's known defect list.
Two reference clients in examples/ are written against this
document and share no code with PlugRL.
| Layer | Choice |
|---|---|
| Transport | WebSocket (RFC 6455), unencrypted ws:// |
| Framing | binary frames only; one message per frame |
| Serialization | msgpack, with two conventions for numeric arrays |
| Schema | four message types, distinguished by a message_type string |
There is no version negotiation, no handshake beyond WebSocket's own, and no extension types. A client that can open a socket and parse msgpack is seven-eighths of the way there; the remaining eighth is the ndarray convention in section 3.
Both endpoints MUST be configured with:
compression = None.permessage-deflatemust not be negotiated. Observations are mostly already-compressed image bytes, and deflating them costs CPU on the hot path for nothing.max_size = None. No frame size limit. A single observation carrying two camera frames is around 180 KB; batched observations and unbounded info dicts pass the 1 MiB library default easily.
The server listens on 0.0.0.0:8000 by default.
The client MAY send an Authorization: Api-Key <key> request header
during the WebSocket handshake.
Gap — the header is not checked.
WebSocketEnvClientAgentsends the header whenapi_keyis set, and no server in this project reads it. The protocol has no authentication. Do not expose a training server to a network you do not control.
Every application message is a msgpack map whose "message_type" key
(a msgpack str) holds one of exactly four lowercase values:
| Value | Direction | Reply |
|---|---|---|
"metadata" |
server to client | none |
"infer" |
client to server | "action" |
"action" |
server to client | "feedback" |
"feedback" |
client to server | none |
These four strings are the protocol's only magic constants besides the two
close reasons in section 7. They are pinned by
tests/test_wire_format.py.
An endpoint receiving a message whose message_type is not the one it
expects MUST treat it as a protocol error (section 7.2). The server does
exactly this; it does not attempt to recover in place.
msgpack maps, arrays, strings, integers, floats, booleans and nil map
directly. Map keys carrying schema names ("message_type", "data",
"env_indices", and so on) are msgpack str.
A multidimensional numeric array travels as a map with binary keys:
{
b"__ndarray__": true,
b"data": <bin>, # raw C-contiguous element bytes
b"dtype": <str>, # a numpy typestr, e.g. "<f4"
b"shape": [d0, d1, ...], # msgpack array of non-negative integers
}
The keys are msgpack bin, not str. In Python that means unpacking
with raw=False yields bytes keys; a decoder that only looks for the text
"__ndarray__" will not find it. This is inherited from openpi (section 9)
and kept for byte compatibility.
data MUST be exactly prod(shape) * itemsize bytes in C (row-major)
order. There is no stride, offset or Fortran-order field.
A zero-dimensional numpy scalar travels differently:
{ b"__npgeneric__": true, b"data": <scalar>, b"dtype": <str> }
where data is an ordinary msgpack number or boolean. Env clients rarely
produce these; decoders should still handle them, since anything a policy
puts in its response may take this form.
dtype is a numpy typestr: a byte-order character, a kind
character, and a decimal item size in bytes.
| Byte order | Meaning |
|---|---|
< |
little-endian |
> |
big-endian |
| |
not applicable (single-byte types) |
| Kind | Meaning | Seen on the wire as |
|---|---|---|
b |
boolean | |b1 — one byte per element, 0 or 1 |
u |
unsigned int | |u1 for images; <u2, <u4, <u8 |
i |
signed int | <i8 for env_indices and step_ids |
f |
float | <f4 for rewards and actions, <f8 for states |
U |
unicode | <U<n> — see section 3.4 |
Kinds V (structured), O (object) and c (complex) are rejected at
pack time and never appear.
A ten-line hand-written parser is enough; both reference clients contain one. This is the only piece of numpy vocabulary a non-Python client must learn — with one exception.
Observation.text is a numpy array of unicode strings, so it is packed by
the same ndarray path as everything else, and its typestr is <U<n> where
n is the character count of the longest string in the array. Each
element occupies 4n bytes: UTF-32 in the stated byte order, NUL-padded
to n code points.
Measured, for ["pick up the red block", "stack"]:
dtype "<U21" shape [2] data 168 bytes
first 8 bytes: 70 00 00 00 69 00 00 00 # 'p', 'i' as UTF-32LE
n is a property of the message, not of the schema — it changes whenever
the longest prompt changes. A non-Python client that wants to read a text
field must therefore implement fixed-width UTF-32 with NUL stripping.
Gap — clients disagree about how to write it.
plugrl-env-clientsendstextas a<Uarray, as above.examples/raw_client.pysends it as a plain msgpack array of strings. Both are accepted, because the server does not validate observation contents — it forwards them to the policy, and whether the policy copes depends on the policy. Of the policies inplugrl-server, onlydummy_policy(which takeslen(obs["text"])) and the openpi transform (which renamestexttoprompt) touch the field, and both accept either form.A future revision should require the msgpack-string-array form and have the server normalise on receipt. Until then, a client that must interoperate with the Python one should send the
<Uarray.
client server
|------------ WebSocket handshake ---------->|
|<---------------- metadata -----------------| exactly once, server speaks first
| |
|------------------ infer ------------------>| \
|<----------------- action ------------------| | repeats until close
| | |
| ... zero or more local environment steps ... |
| | |
|----------------- feedback ---------------->| /
| (no reply) |
On a given connection, application messages strictly alternate:
metadata ( infer action feedback )*
The server's connection handler is a straight-line loop that blocks on
recv() for an infer, replies with an action, then blocks on recv()
again for exactly one feedback. It has no queue and no dispatcher. A
client that sends two infer messages without a feedback between them
will have its second infer parsed as a feedback, fail validation, and be
disconnected with a resync request (section 7.2).
A feedback may arrive many environment steps after the action it
answers. With an action horizon of H, the client executes up to H
local steps before the chunk is exhausted and feedback is due. The server
waits, up to FEEDBACK_WAIT_TIMEOUT = 60 s (section 7.3).
This is the part an independent implementation is most likely to get wrong.
infer, action and feedback each carry a set of environment indices.
Within one infer/action/feedback cycle:
- the
action'senv_idsMUST equal theinfer'senv_indices, in the same order — the server asserts this; - the
feedback'senv_indicesneed not equal either of them, and in a multi-environment client routinely does not.
The two sets answer different questions. infer carries the environments
whose action chunk has just run out; feedback carries the environments
whose chunk finished on this environment step. Once environments
desynchronise — the first time one of them terminates early — those are
different sets.
The pairing between an action and the following feedback is therefore
pure flow control, not a semantic association. The server routes
feedback by env index into per-environment state, not by position in the
stream. An implementer who assumes "the feedback tells me about the actions
I just received" will be wrong as soon as a second environment exists.
The alternation still holds exactly, because an environment can only re-enter the "needs a chunk" set at the top of the iteration after the one in which it entered the "chunk finished" set. One feedback per action, every time.
env_indices are small non-negative integers naming environments within
one connection. The server keeps its per-environment maps inside the
connection handler, so index 0 from one client and index 0 from another
are unrelated. Clients running several processes against one server do not
need to coordinate index ranges.
Shapes are written with n = number of environments in this message,
H = action horizon, da = the environment's action shape (usually a
one-tuple).
{ "message_type": "metadata",
"data": <map> }
Sent once, immediately after the handshake, before the client sends
anything. A client MUST read it before its first infer.
data is a map of descriptive keys. A client MUST NOT require any of
them, and MUST ignore keys it does not recognise. What the server puts
there today:
| Key | Meaning |
|---|---|
protocol_version |
the version of this document the server implements — 1 |
server |
"plugrl-server" |
server_version |
the package version |
algorithm |
the learning algorithm's class name |
policy |
the policy's class name |
action_horizon |
H, the number of steps in an action chunk |
action_dim |
the width of one action |
features |
the optional features of section 10.1 the server implements, as a list of strings |
action_horizon and action_dim are read off the policy, which is free
not to declare them. A key that is absent means the server does not know,
never that the value is zero or a default — so a client that needs the
shape and does not find it must be configured with it, exactly as before.
That rule is what makes the rest of the map trustworthy.
The server merges anything its caller passes over the top, so an operator can add a run identifier or correct a value the introspection got wrong.
Historical note. Through 2026-09-10 this message was always
{}— every entry point defaulted it and nothing filled it in, so the action shape had to travel out of band. A client written against that behaviour still works, which is why none of these keys are required.
{ "message_type": "infer",
"data": <observation>,
"env_indices": <ndarray "<i8" shape [n]>,
"step_ids": <ndarray "<i8" shape [n]> }
The observation is a three-key map, batched along a leading axis of n:
{ "images": { <name>: <ndarray "|u1" shape [n, h, w, c]>, ... },
"states": { <name>: <ndarray "<f8" shape [n, d]>, ... },
"text": <see section 3.4, shape [n]> }
images and states may each be empty; camera names and state names are
environment-defined and the server does not check them. Every array under
both keys MUST share the same leading dimension n.
n MAY differ between successive infer messages on one connection.
This is the normal case: the batch is whichever environments happen to need
a chunk, and it is ragged by construction.
An infer may also carry a reuse key, when the server offers the
reuse-feedback-obs feature of section 10.1. A row it marks carries no
observation: the server takes it from that environment's last feedback.
Without the feature the key is not sent.
Gap —
step_idsis accepted and ignored. The client sends, per environment, the number of completed action chunks in the current episode, reset to0when the episode ends. Neither server reads it. It is validated for presence only; a client must send it, and may send zeros.
{ "message_type": "action",
"data": { "env_ids": <ndarray "<i8" shape [n]>,
"action": <ndarray shape [H, n, *da]> } }
The action array is time-major: horizon first, environment second. The
server produces [n, H, *da] internally and transposes on the way out.
The client MAY consume any prefix of the horizon and re-plan early; it
MUST NOT ask for more steps than H. The dtype is the environment's
action dtype and is not renegotiated.
Gap — the two servers disagree about this message.
WebSocketAgentServersendsenv_idsalongsideaction, as above.RayAgentServersends{"action": ...}with noenv_idskey at all, ignores the request'senv_indices, keeps one environment's worth of state per connection, and buffers the action chunk server-side — delivering one step perinferwhen the algorithm setsbreak_action_chunk. It is an earlier single-environment dialect that was not updated when multi-environment support landed.A client conforming to this document should read
env_idswhen present and otherwise assume the request's ownenv_indices. Bringing the Ray server to this specification, or retiring it, is tracked work rather than a protocol choice.
{ "message_type": "feedback",
"env_indices": <ndarray "<i8" shape [m]>,
"step_ids": <ndarray "<i8" shape [m]>,
"data": {
"obs": <observation, batched to m>,
"rewards": <ndarray, float kind, shape [m]>,
"terminated": <ndarray "|b1" shape [m]>,
"truncated": <ndarray "|b1" shape [m]>,
"info": <map> } }
There is no reply. The next thing on the wire is the client's next
infer.
obs is the observation after the last action of the chunk — the
environment's next state, not the one that was sent in infer.
rewards is the sum over the chunk, not the last step's reward. The
client accumulates reward across every environment step of the chunk and
flushes the total when the chunk ends. Credit assignment reaching the
learner is therefore at chunk granularity. A client that reports only the
final step's reward will silently train on a different MDP.
plugrl-env-client sends <f4. The server does not inspect the width, so a
<f8 reward is accepted; float32 is the convention rather than a rule, and
examples/conformance_server.py reports the difference as a note rather
than a violation.
terminated and truncated are the flags from the final step of the
chunk, carrying Gymnasium's usual distinction: terminated means the
episode ended by the environment's own rules, truncated means it was cut
short (time limit, external stop). A chunk is also flushed early when either
flag is set, so a terminal transition is never buried inside a chunk.
info is a free-form map. {} is valid and is what both reference clients
send. Anything else needs care, and the rule is narrower than it looks. What
the servers read from it is in the next subsection.
Gap - a non-empty
infocan close the connection. The server unbatches it by looking for the first value that is an ndarray - at the top level, or one level into a nested map - and takingmfrom that array's leading axis, then slicing every ndarray value by it; a scalar or string sitting alongside is broadcast to allmenvironments, which is the documented behaviour. But if no value is an ndarray, the map is returned whole as a single per-environment entry regardless ofm
{"task": "pick"}withm= 2 yields one entry, and so does a correctly batched msgpack list of lengthm, because a list is not an ndarray (unbatch_aggregateinplugrl-server'scommon/data_utils.py).What follows changed with plugrl-server #108. The server now checks that the observations, rewards, both flags and a non-empty
infoeach holdmentries. A mismatch is a protocol error: it closes that connection with1001andplugrl-server-resync, as for any malformed message, and its other clients go on. A client that sends the sameinfoagain after reconnecting is closed again.Before #108 the check was an
assertin the connection handler. Any exception there is fatal to the server, so the connection closed with 1011 andInternal server error., and the server then stopped, for every client. This paragraph used to mention only the closed connection. plugrl-server's docs audit found the rest on 2026-09-29.So a non-empty
infois safe only when at least one of its values, or a value one level into a nested map, is an ndarray of lengthm, or whenmis 1, where the mismatch cannot arise. The previous wording of this paragraph - that a value which is notm-shaped is passed through to every environment unchanged - is true only in the first of those cases. This has not bitten anyone because the reference clients send{}.
One key, episode, and only on a transition whose terminated or
truncated is set. It carries the finished episode's statistics, each
batched to m like any other value:
"episode": {
"r": <ndarray, float kind, shape [m]>, the episode's return
"l": <ndarray, int kind, shape [m]>, its length in environment steps
"s": <ndarray "|b1" shape [m]>, whether it succeeded
"mask": <ndarray "|b1" shape [m]> } whether this entry is a finished episode
The training algorithms (ppo, fpo, dppo) average the episodes recorded
since the last learn step into rollout/reward, rollout/length and
rollout/success, and start again after it. evaluation also logs each
episode as it finishes. An entry whose mask is false is skipped, and a
missing mask counts as true. Nothing else in info is read.
Leaving episode out is valid and changes nothing about training: the same
transitions are stored and learned from. Only those metrics read 0. That is
how it was found: a C++ client written from this document trained Pendulum
normally while its logged return stayed at 0, until it sent episode
(plugrl-server E44).
plugrl-env-client sends more than this, and the servers read none of the
rest:
episodeis full-length. Every environment in the batch has an entry;maskis true only where the episode ended at this step, and the other entries are zeros.episodehas a fifth key,mean_success_rate: a float over the client's last 100 episodes.- A top-level
is_step_successof shape[m]rides on every feedback.
When the server looks for m in info, it also looks one level into a
nested map. So {"episode": {...}} on its own is unbatched correctly, even
though none of its top-level values is an ndarray.
The client MUST send the observation that the terminal step returned,
not the first observation of the next episode. This corresponds to
Gymnasium's AutoresetMode.NEXT_STEP: the step reporting done returns the
terminal observation, and the reset happens on the following call.
An environment that resets inside its own step() violates this silently —
nothing raises, and the terminal transition simply carries the wrong
observation into the learner's buffer. plugrl-env-client pins the declared
mode of every environment it ships and demonstrates the corruption
executably in its tests/test_terminal_observation.py.
Not part of the wire format, but it determines what a client sees.
The server does not answer an infer immediately. It queues the request and
a scheduler drains the queue into one batch, running the policy once for all
of them. A request therefore waits for two things: the scheduler's polling
interval (SCHEDULER_SLEEP_INTERVAL, 0.1 ms) and the batch-readiness
condition — by default one queued request per connected client, or
mini_infer_batch_size environments when that is set.
Two consequences for a client:
- Latency is not independent of other clients. A client alone on a
server is answered as fast as the scheduler polls. A client sharing one
with slower peers waits for them, up to
INFER_READY_TIMEOUT= 5 s, after which the server logs a warning and keeps waiting. - A client must not assume its request is the whole batch. Actions come
back sliced out of a larger tensor; the slice boundaries are the request's
own
n.
When the algorithm signals that training is finished, the server closes with close code 1001 (going away) and reason:
plugrl-server-stop
A client seeing this reason should exit rather than reconnect, unless it was explicitly configured to wait for the next run.
When a message fails validation — wrong message_type, missing required
key — the server closes with 1001 (going away) and reason:
plugrl-server-resync
This means "your stream and mine are out of step". A client should
reconnect, read the fresh metadata, and resume from a new infer. Any
feedback it was holding is stale and MUST be dropped: the server's
per-environment state went with the closed connection.
If 60 s pass after an action without a feedback, the server closes with
1001 and reason Feedback timeout. Environments slower than that must
be split across more connections, or the constant raised on both sides.
An unhandled server-side exception closes with 1011 (internal error) and
reason Internal server error.. Unlike openpi (section 9), PlugRL does
not send the traceback as a text frame first.
A client should nonetheless be prepared to receive a text frame where it expected binary, and treat it as a fatal server-side error rather than attempting to unpack it.
Everything above is the server closing. A client may also stop first - it has collected the episodes it was asked for, or its operator interrupted it - and that is not an error. The server keeps no state that outlives the connection, so a client that is done costs nothing beyond the feedback it had not yet sent. A client that intends to come back pays more than that; see section 7.6.
A client that is finished SHOULD send a WebSocket close frame with status
1000 (normal closure) before dropping the socket, as RFC 6455 section
5.5.1 asks. Nothing breaks without one - the server sees the connection end
either way - but the difference is visible in its logs, as
ConnectionClosedError: no close frame received or sent rather than a clean
ConnectionClosedOK, and an operator reading those logs should not have to
wonder whether a client crashed.
examples/conformance_server.py reports a missing close frame as a note, not
a violation, which is the level this rule deserves.
Historical note. Until 2026-09-11 the C++ reference client did not send one. It went unnoticed because in every test until then the server ran out of steps first and closed the connection itself; E7, where the client finishes first, made it visible on all 45 runs.
Section 7.2 says the server's per-environment state went with the closed connection. That is true of every close, not only a resync, and it is the one thing a reconnecting client has to reason about.
The server holds, per environment and for exactly one connection: the previous observation, the policy step state, and the terminated / truncated flags. A new connection begins with none of them. So:
- a client that reconnects MUST drop any
feedbackit was holding; - it reads a fresh
metadataand resumes from a freshinfer; - a
feedbackwhoseactionarrived on an earlier connection describes a transition the server can no longer complete, because it has no previous observation to attach it to. Sending it produces a transition built from nothing, which is worse than the lost step it was trying to save.
A server SHOULD say something when it receives feedback for an environment it has no step state for. That condition has exactly one cause - the connection was replaced mid-run - and storing the transition silently puts a hole in the training data that nothing downstream can detect.
Gap - the reference client breaks this on one path.
plugrl-env-clientrun with--reconnect-on-server-stopresends thefeedbackit was holding instead of dropping it, when the close it hit was the server'splugrl-server-stop. Inwebsocket_env_client_agent.py,feedback()'sSERVER_STOP_REASONbranchcontinues its retry loop, which opens a new connection, reads fresh metadata, and re-sends the same payload - the exact resend the historical note below says this section exists to prevent. Every other close path in that method returns and drops the transition, andtests/test_reconnect_drops_feedback.pycovers three of them (a keepalive timeout, a plain 1000 close, a resync) but not this one. It is the only path in the reference client that violates the MUST above.
Where reconnects come from. Nothing in this protocol causes them and nothing in it can prevent them: a suspended laptop, a flaky link, an operator restarting the server. The rule above is written in terms of the reconnect rather than its cause, because the cost is the same either way, and because a client cannot tell the causes apart from where it sits.
Historical note. This section exists because of a run that dropped its connection after the machine it was on suspended for nearly two hours. The reconnect was handled, and the transition that crossed it was not: the env client resent the
feedbackit was holding, and the server completed it from an empty observation and stored it. The first diagnosis blamed WebSocket keepalive pings and a long learn step, which measurement then ruled out -plugrl-serverrunslearnoff the event loop, and learns of 190 s produce no ping timeout. The cause was mundane. The hole in the protocol was not, and had nothing to do with it.
A client conforms to version 1 if it:
- connects over WebSocket with compression disabled and no frame size cap;
- reads one
metadatamessage before sending anything, and requires no particular key in it; - emits binary frames containing msgpack maps with a
message_typestr; - encodes arrays with bin keys
__ndarray__/data/dtype/shape, C-contiguous, and parses the typestr rather than assuming a dtype; - sends exactly one
feedbackfor everyactionreceived, and never twoinfermessages in a row; - tolerates its
feedbackenv set differing from itsinferenv set; - reads
actionas[H, n, *da], time-major, and readsenv_idswhen present; - reports chunk-summed reward, and the terminal observation on the step that reports done;
- handles close reasons
plugrl-server-stopandplugrl-server-resyncdifferently; - drops any held
feedbackwhen a connection closes, and never sendsfeedbackfor anactionthat arrived on an earlier connection; - treats a text frame as a fatal error;
- uses a feature of section 10 only when the server lists it, and, for
reuse-feedback-obs, reuses only an observation the server holds.
plugrl-conformance (examples/conformance_server.py is a shim for it) has
two modes.
By default it watches one connection. It checks what that connection
shows: framing, alternation, env indices, observation shape, and the
feedback payload's keys, dtypes and lengths. It reports what it accepts but
cannot require as a note rather than a failure. Both reference clients pass
it with one note: they send text as a msgpack string array rather than a
<U array, which is the section 3.4 Gap. An unexercised clause leaves no
trace in its report, and from one connection to an unknown environment it
cannot tell what the client did with an action.
With --probe the client runs the probe environment of section 8.1.
From each feedback the checker then works out how many steps of the chunk
the client ran and which action it applied last, and it drives the
connection instead of watching it. In that mode it checks every clause of
the list above except these:
- reading
env_ids: section 4.3 requires them to equal theinfer'senv_indices, so a client that ignores them behaves identically; - tolerating a
feedbackenv set that differs from theinferset: the client chooses its sets, and the server cannot make them differ; - requiring no particular
metadatakey: the checker sends all of them; - dropping held
feedbackafter a close other than a resync.
--scenario selects how the connection is driven:
basic: the checker waits 0.3 s before speaking, so a client that sends first is caught. Itsmetadatacarries an unknown 2 MiB key, so a client that kept its library's 1 MiB frame cap cannot receive it, and one that rejects unknown keys fails. It checks that the handshake did not offerpermessage-deflate. It answers everyinferwith float64, time-major actions whose values encode their position. Once the exchanges are done it closes withplugrl-server-stop, after which the client must exit 0 and not reconnect.resync: asbasic, but halfway through it closes withplugrl-server-resyncright after sending anaction, so the client is holding afeedback. The client must reconnect, and its first message on the new connection must be aninfer.text: it answers the firstinferwith a text frame, which the client must treat as fatal: it closes the connection and sends nothing more.all: each in turn, starting the client once per scenario.
Both reference clients pass all three with the same note:
examples/raw_client.py --probe and plugrl_client --probe, the C++ one,
which CI checks on every change. The Python client's --bug option breaks
one clause at a time, and tests/test_conformance_probe.py checks that the
checker names each one.
A client that wants its handling of actions checked, and not only its
messages, runs this environment for plugrl-conformance --probe. It needs
no simulator, and is a few lines in any language.
-
The client runs
nprobe envs, with indices0ton - 1as itsenv_indices. Indices stay below 1000. -
Env
i's episode lastsL_i = 3 + 2 * (i mod 3)steps: 3, 5, 7, 3, 5, ... So within one chunk, different envs end their episodes at different steps. -
Its observation has no images and two states:
t: float64[n, 1], the number of steps taken in the current episode,0after a reset;a: float64[n, d], the action the env applied on its last step, wheredis the action width. Its value after a reset is not checked, and before the client knows the width it may haved = 1.
textmay be anything. -
A step applies the action, adds 1 to
t, stores the action ina, and pays a reward of 1. It reportsterminatedwhentreachesL_i. The probe env never truncates. -
A reset sets
tback to0.
The checker sends action values that encode their position: element
[k, row, d] for env j is 10000 k + 10 j + d. A feedback therefore says
exactly what the client did with the chunk. t minus the t of the
infer it answered is the number of steps it ran, which must be between 1
and H. With a reward of 1 per step, the reward must equal that number,
which is the chunk-sum rule. a must be the chunk's action at the last step
run, for that env, which checks time-major order, the step order, and the
dtype. A terminated env must report t = L_i, its own terminal
observation, and its next infer must report t = 0.
The checklist above is for clients. A training server conforms if it:
- sends one
metadatamessage, before anything else, on every connection; - answers each
inferwith anactionthat is time-major,[H, n, *da], withenv_idsequal to theinfer'senv_indices, in order, and that agrees with theaction_horizonandaction_dimit declared; - accepts what a client may send: a
feedbackenv set unlike theinfer's, annthat varies, and frames larger than 1 MiB; - closes a connection with 1001 and
plugrl-server-resyncon a malformed message, twoinfers in a row, or aninfoit cannot split per environment, and stays up for its other clients; - keeps env indices connection-scoped;
- ends a run with 1001 and
plugrl-server-stop; - if it lists
reuse-feedback-obs, answers aninferthat reuses observations, and closes with 1001 andplugrl-server-resyncon one that reuses an observation it does not hold.
plugrl-conformance-server checks each of these from the client's side of
the wire:
plugrl-conformance-server --port 8000 --state-dim 3 --until-stopIt sends a states-only observation, states["obs"] of --state-dim values
(--state-key and --image-key change that), so it can drive any server
whose policy reads one. examples/reference_server.py is a server written
against this document that trains nothing and passes every check. Its
--bug option breaks one clause at a time, and
tests/test_server_conformance.py checks that the grader names each one.
plugrl-server passes too. Its version from before plugrl-server#108 fails
two checks: an info it could not split closed that connection with 1011,
and the server then refused new connections.
What the server does with a frame is not visible from the wire, so section 7.6's SHOULD - saying something about feedback for an environment it holds no step state for - is not checked.
PlugRL's serialization is openpi's. msgpack_numpy.py is taken from
openpi (Apache-2.0,
attributed in the file header and in NOTICE), reformatted only. The
array encoding is byte-identical, deliberately, so that the hard part of
writing a client is shared between the two ecosystems.
The message layer differs:
| openpi | PlugRL | |
|---|---|---|
| First server message | the metadata map, bare | {"message_type": "metadata", "data": ...} |
| Request | the observation map, bare | envelope + env_indices + step_ids |
| Response | the action map, bare, plus server_timing |
envelope + env_ids + action |
| Return channel | none | feedback: reward, done flags, next observation |
| Batching | one observation per request | ragged multi-environment batch |
| Errors | traceback sent as a text frame, then close 1011 | close 1011, no traceback frame |
| Health check | GET /healthz |
none |
PlugRL is openpi's serving protocol plus a feedback return channel. That is the whole difference, and it is the difference between serving a policy and training one. An openpi client cannot drive a PlugRL server unmodified, and vice versa, but the array codec — the part that takes real work in C++ or Rust — carries across unchanged.
This is version 1. The server publishes it as protocol_version in the
metadata message (section 5.1). A client MUST NOT require the key, and
the server does not adapt to a client's version.
What is negotiated is features. A feature is an addition to version 1
that a client may use and a server may offer, and nothing else changes for
either side when it is absent. The server lists the ones it implements in
metadata's features. A client uses a feature only if it is listed there,
and a server that lists one still accepts every client that does not use it.
So a new client works with an old server, and an old client with a new one.
A change that cannot be made this way is breaking, and moves
protocol_version to 2.
Correction, 2026-09-11. This paragraph used to say version 1 "has no version field on the wire" and attributed to section 5.1 a remark that the metadata message was the obvious place for one. Both halves went stale on 2026-09-10, when the metadata message gained contents:
protocol_versionhas been in it on every connection since, inplugrl-serverand inexamples/conformance_server.pyalike, and section 5.1 lists it as a key the server sends rather than as a suggestion. What survives is the absence of negotiation, which is what the sentence was reaching for.
Changing any of the four message_type strings, the two close reasons, the
ndarray key names, the action array's axis order, or the meaning of
rewards is a breaking change.
The first three are pinned in this repository, and a change to them fails a
test rather than a deployment: tests/test_wire_format.py holds the four
message_type strings and the two close reasons,
tests/test_spec_conformance.py the four ndarray key names. The last two
are not, and cannot be - this is the codec repository, and the axis order
and the chunk-sum rule are properties of how the two sides behave, not of
what the packer emits. Transposing the action chunk or redefining rewards
as the last step's reward passes every codec test here. They are checked
instead by plugrl-conformance --probe (section 8.1), against any client.
Without it, every observation crosses the link twice. The client sends it in
the feedback that ends a chunk, as the step's next observation, and then
sends the same observation again in the next infer, as the state the new
chunk starts from. Only after a reset do the two differ. That second copy is
the factor of two in plugrl-server's E43 cost model: an exchange costs a
fixed latency plus twice the observation's bytes over the link.
With the feature, the infer may say "the observation you already have":
{ "message_type": "infer",
"data": <observation, batched to the rows that are not reused>,
"env_indices": <ndarray "<i8" shape [n]>,
"step_ids": <ndarray "<i8" shape [n]>,
"reuse": <ndarray "|b1" shape [n]> }
reuse[i]true means envenv_indices[i]sends no observation. Its observation is theobsrow of the lastfeedbackthat named it on this connection.datais batched to the rows wherereuseis false, in theirenv_indicesorder. With every row reused, its arrays have a leading dimension of 0.- A client MUST send
reusefalse for an env whose last feedback on this connection reportedterminatedortruncated, since that feedback carried the terminal observation and not the reset one. It MUST also send it false for an env with no feedback on this connection yet, which includes every env just after a reconnect. - A server that lists the feature MUST close a connection whose
inferreuses an observation it does not hold - an env never fed back on this connection, or one whose last feedback ended its episode - with 1001 andplugrl-server-resync. - An
inferwithoutreusesends every observation, as in version 1.
Both checkers know the feature. plugrl-conformance --features reuse-feedback-obs offers it to a client and checks how the client uses it.
plugrl-conformance-server exercises it when the server lists it.