Skip to content

Latest commit

 

History

History
899 lines (717 loc) · 40.1 KB

File metadata and controls

899 lines (717 loc) · 40.1 KB

The PlugRL wire protocol

Version 1 · written 2026-09-10 against plugrl-protocol@22b324c, plugrl-server@06e9f63 and plugrl-env-client@b9edc5b.

This document specifies the bytes exchanged between a training server, which holds the policy and the learning algorithm, and an env client, which runs environments and asks for actions. Anything able to speak WebSocket and msgpack can be an env client: a Python process, a ROS node, a robot's onboard C++ controller.

The specification is written from the implementation, not the other way round. Where the implementation is under-specified, ambiguous, or where two of its own servers disagree, this document says so in a box marked Gap rather than inventing a rule. Those boxes are the honest part; treat them as the protocol's known defect list.

Two reference clients in examples/ are written against this document and share no code with PlugRL.


1. Layers

Layer Choice
Transport WebSocket (RFC 6455), unencrypted ws://
Framing binary frames only; one message per frame
Serialization msgpack, with two conventions for numeric arrays
Schema four message types, distinguished by a message_type string

There is no version negotiation, no handshake beyond WebSocket's own, and no extension types. A client that can open a socket and parse msgpack is seven-eighths of the way there; the remaining eighth is the ndarray convention in section 3.

1.1 Connection options

Both endpoints MUST be configured with:

  • compression = None. permessage-deflate must not be negotiated. Observations are mostly already-compressed image bytes, and deflating them costs CPU on the hot path for nothing.
  • max_size = None. No frame size limit. A single observation carrying two camera frames is around 180 KB; batched observations and unbounded info dicts pass the 1 MiB library default easily.

The server listens on 0.0.0.0:8000 by default.

1.2 Authentication

The client MAY send an Authorization: Api-Key <key> request header during the WebSocket handshake.

Gap — the header is not checked. WebSocketEnvClientAgent sends the header when api_key is set, and no server in this project reads it. The protocol has no authentication. Do not expose a training server to a network you do not control.


2. Message envelope

Every application message is a msgpack map whose "message_type" key (a msgpack str) holds one of exactly four lowercase values:

Value Direction Reply
"metadata" server to client none
"infer" client to server "action"
"action" server to client "feedback"
"feedback" client to server none

These four strings are the protocol's only magic constants besides the two close reasons in section 7. They are pinned by tests/test_wire_format.py.

An endpoint receiving a message whose message_type is not the one it expects MUST treat it as a protocol error (section 7.2). The server does exactly this; it does not attempt to recover in place.


3. Serialization

3.1 Ordinary values

msgpack maps, arrays, strings, integers, floats, booleans and nil map directly. Map keys carrying schema names ("message_type", "data", "env_indices", and so on) are msgpack str.

3.2 Arrays

A multidimensional numeric array travels as a map with binary keys:

{
  b"__ndarray__": true,
  b"data":  <bin>,          # raw C-contiguous element bytes
  b"dtype": <str>,          # a numpy typestr, e.g. "<f4"
  b"shape": [d0, d1, ...],  # msgpack array of non-negative integers
}

The keys are msgpack bin, not str. In Python that means unpacking with raw=False yields bytes keys; a decoder that only looks for the text "__ndarray__" will not find it. This is inherited from openpi (section 9) and kept for byte compatibility.

data MUST be exactly prod(shape) * itemsize bytes in C (row-major) order. There is no stride, offset or Fortran-order field.

A zero-dimensional numpy scalar travels differently:

{ b"__npgeneric__": true, b"data": <scalar>, b"dtype": <str> }

where data is an ordinary msgpack number or boolean. Env clients rarely produce these; decoders should still handle them, since anything a policy puts in its response may take this form.

3.3 The typestr

dtype is a numpy typestr: a byte-order character, a kind character, and a decimal item size in bytes.

Byte order Meaning
< little-endian
> big-endian
| not applicable (single-byte types)
Kind Meaning Seen on the wire as
b boolean |b1 — one byte per element, 0 or 1
u unsigned int |u1 for images; <u2, <u4, <u8
i signed int <i8 for env_indices and step_ids
f float <f4 for rewards and actions, <f8 for states
U unicode <U<n> — see section 3.4

Kinds V (structured), O (object) and c (complex) are rejected at pack time and never appear.

A ten-line hand-written parser is enough; both reference clients contain one. This is the only piece of numpy vocabulary a non-Python client must learn — with one exception.

3.4 Text is the one real wart

Observation.text is a numpy array of unicode strings, so it is packed by the same ndarray path as everything else, and its typestr is <U<n> where n is the character count of the longest string in the array. Each element occupies 4n bytes: UTF-32 in the stated byte order, NUL-padded to n code points.

Measured, for ["pick up the red block", "stack"]:

dtype  "<U21"        shape [2]        data 168 bytes
first 8 bytes: 70 00 00 00 69 00 00 00     # 'p', 'i' as UTF-32LE

n is a property of the message, not of the schema — it changes whenever the longest prompt changes. A non-Python client that wants to read a text field must therefore implement fixed-width UTF-32 with NUL stripping.

Gap — clients disagree about how to write it. plugrl-env-client sends text as a <U array, as above. examples/raw_client.py sends it as a plain msgpack array of strings. Both are accepted, because the server does not validate observation contents — it forwards them to the policy, and whether the policy copes depends on the policy. Of the policies in plugrl-server, only dummy_policy (which takes len(obs["text"])) and the openpi transform (which renames text to prompt) touch the field, and both accept either form.

A future revision should require the msgpack-string-array form and have the server normalise on receipt. Until then, a client that must interoperate with the Python one should send the <U array.


4. The exchange

4.1 Sequence

client                                     server
  |------------ WebSocket handshake ---------->|
  |<---------------- metadata -----------------|   exactly once, server speaks first
  |                                            |
  |------------------ infer ------------------>|   \
  |<----------------- action ------------------|    |  repeats until close
  |                                            |    |
  |   ... zero or more local environment steps ...   |
  |                                            |    |
  |----------------- feedback ---------------->|   /
  |                (no reply)                  |

4.2 The alternation invariant

On a given connection, application messages strictly alternate:

metadata  ( infer  action  feedback )*

The server's connection handler is a straight-line loop that blocks on recv() for an infer, replies with an action, then blocks on recv() again for exactly one feedback. It has no queue and no dispatcher. A client that sends two infer messages without a feedback between them will have its second infer parsed as a feedback, fail validation, and be disconnected with a resync request (section 7.2).

A feedback may arrive many environment steps after the action it answers. With an action horizon of H, the client executes up to H local steps before the chunk is exhausted and feedback is due. The server waits, up to FEEDBACK_WAIT_TIMEOUT = 60 s (section 7.3).

4.3 The env sets in one cycle need not match

This is the part an independent implementation is most likely to get wrong.

infer, action and feedback each carry a set of environment indices. Within one infer/action/feedback cycle:

  • the action's env_ids MUST equal the infer's env_indices, in the same order — the server asserts this;
  • the feedback's env_indices need not equal either of them, and in a multi-environment client routinely does not.

The two sets answer different questions. infer carries the environments whose action chunk has just run out; feedback carries the environments whose chunk finished on this environment step. Once environments desynchronise — the first time one of them terminates early — those are different sets.

The pairing between an action and the following feedback is therefore pure flow control, not a semantic association. The server routes feedback by env index into per-environment state, not by position in the stream. An implementer who assumes "the feedback tells me about the actions I just received" will be wrong as soon as a second environment exists.

The alternation still holds exactly, because an environment can only re-enter the "needs a chunk" set at the top of the iteration after the one in which it entered the "chunk finished" set. One feedback per action, every time.

4.4 Environment indices are connection-scoped

env_indices are small non-negative integers naming environments within one connection. The server keeps its per-environment maps inside the connection handler, so index 0 from one client and index 0 from another are unrelated. Clients running several processes against one server do not need to coordinate index ranges.


5. Message schemas

Shapes are written with n = number of environments in this message, H = action horizon, da = the environment's action shape (usually a one-tuple).

5.1 metadata — server to client

{ "message_type": "metadata",
  "data": <map> }

Sent once, immediately after the handshake, before the client sends anything. A client MUST read it before its first infer.

data is a map of descriptive keys. A client MUST NOT require any of them, and MUST ignore keys it does not recognise. What the server puts there today:

Key Meaning
protocol_version the version of this document the server implements — 1
server "plugrl-server"
server_version the package version
algorithm the learning algorithm's class name
policy the policy's class name
action_horizon H, the number of steps in an action chunk
action_dim the width of one action
features the optional features of section 10.1 the server implements, as a list of strings

action_horizon and action_dim are read off the policy, which is free not to declare them. A key that is absent means the server does not know, never that the value is zero or a default — so a client that needs the shape and does not find it must be configured with it, exactly as before. That rule is what makes the rest of the map trustworthy.

The server merges anything its caller passes over the top, so an operator can add a run identifier or correct a value the introspection got wrong.

Historical note. Through 2026-09-10 this message was always {} — every entry point defaulted it and nothing filled it in, so the action shape had to travel out of band. A client written against that behaviour still works, which is why none of these keys are required.

5.2 infer — client to server

{ "message_type": "infer",
  "data":        <observation>,
  "env_indices": <ndarray "<i8" shape [n]>,
  "step_ids":    <ndarray "<i8" shape [n]> }

The observation is a three-key map, batched along a leading axis of n:

{ "images": { <name>: <ndarray "|u1" shape [n, h, w, c]>, ... },
  "states": { <name>: <ndarray "<f8" shape [n, d]>,       ... },
  "text":   <see section 3.4, shape [n]> }

images and states may each be empty; camera names and state names are environment-defined and the server does not check them. Every array under both keys MUST share the same leading dimension n.

n MAY differ between successive infer messages on one connection. This is the normal case: the batch is whichever environments happen to need a chunk, and it is ragged by construction.

An infer may also carry a reuse key, when the server offers the reuse-feedback-obs feature of section 10.1. A row it marks carries no observation: the server takes it from that environment's last feedback. Without the feature the key is not sent.

Gap — step_ids is accepted and ignored. The client sends, per environment, the number of completed action chunks in the current episode, reset to 0 when the episode ends. Neither server reads it. It is validated for presence only; a client must send it, and may send zeros.

5.3 action — server to client

{ "message_type": "action",
  "data": { "env_ids": <ndarray "<i8" shape [n]>,
            "action":  <ndarray shape [H, n, *da]> } }

The action array is time-major: horizon first, environment second. The server produces [n, H, *da] internally and transposes on the way out.

The client MAY consume any prefix of the horizon and re-plan early; it MUST NOT ask for more steps than H. The dtype is the environment's action dtype and is not renegotiated.

Gap — the two servers disagree about this message. WebSocketAgentServer sends env_ids alongside action, as above. RayAgentServer sends {"action": ...} with no env_ids key at all, ignores the request's env_indices, keeps one environment's worth of state per connection, and buffers the action chunk server-side — delivering one step per infer when the algorithm sets break_action_chunk. It is an earlier single-environment dialect that was not updated when multi-environment support landed.

A client conforming to this document should read env_ids when present and otherwise assume the request's own env_indices. Bringing the Ray server to this specification, or retiring it, is tracked work rather than a protocol choice.

5.4 feedback — client to server

{ "message_type": "feedback",
  "env_indices": <ndarray "<i8" shape [m]>,
  "step_ids":    <ndarray "<i8" shape [m]>,
  "data": {
    "obs":        <observation, batched to m>,
    "rewards":    <ndarray, float kind, shape [m]>,
    "terminated": <ndarray "|b1" shape [m]>,
    "truncated":  <ndarray "|b1" shape [m]>,
    "info":       <map> } }

There is no reply. The next thing on the wire is the client's next infer.

obs is the observation after the last action of the chunk — the environment's next state, not the one that was sent in infer.

rewards is the sum over the chunk, not the last step's reward. The client accumulates reward across every environment step of the chunk and flushes the total when the chunk ends. Credit assignment reaching the learner is therefore at chunk granularity. A client that reports only the final step's reward will silently train on a different MDP.

plugrl-env-client sends <f4. The server does not inspect the width, so a <f8 reward is accepted; float32 is the convention rather than a rule, and examples/conformance_server.py reports the difference as a note rather than a violation.

terminated and truncated are the flags from the final step of the chunk, carrying Gymnasium's usual distinction: terminated means the episode ended by the environment's own rules, truncated means it was cut short (time limit, external stop). A chunk is also flushed early when either flag is set, so a terminal transition is never buried inside a chunk.

info is a free-form map. {} is valid and is what both reference clients send. Anything else needs care, and the rule is narrower than it looks. What the servers read from it is in the next subsection.

Gap - a non-empty info can close the connection. The server unbatches it by looking for the first value that is an ndarray - at the top level, or one level into a nested map - and taking m from that array's leading axis, then slicing every ndarray value by it; a scalar or string sitting alongside is broadcast to all m environments, which is the documented behaviour. But if no value is an ndarray, the map is returned whole as a single per-environment entry regardless of m

  • {"task": "pick"} with m = 2 yields one entry, and so does a correctly batched msgpack list of length m, because a list is not an ndarray (unbatch_aggregate in plugrl-server's common/data_utils.py).

What follows changed with plugrl-server #108. The server now checks that the observations, rewards, both flags and a non-empty info each hold m entries. A mismatch is a protocol error: it closes that connection with 1001 and plugrl-server-resync, as for any malformed message, and its other clients go on. A client that sends the same info again after reconnecting is closed again.

Before #108 the check was an assert in the connection handler. Any exception there is fatal to the server, so the connection closed with 1011 and Internal server error., and the server then stopped, for every client. This paragraph used to mention only the closed connection. plugrl-server's docs audit found the rest on 2026-09-29.

So a non-empty info is safe only when at least one of its values, or a value one level into a nested map, is an ndarray of length m, or when m is 1, where the mismatch cannot arise. The previous wording of this paragraph - that a value which is not m-shaped is passed through to every environment unchanged - is true only in the first of those cases. This has not bitten anyone because the reference clients send {}.

What the servers read from info

One key, episode, and only on a transition whose terminated or truncated is set. It carries the finished episode's statistics, each batched to m like any other value:

"episode": {
  "r":    <ndarray, float kind, shape [m]>,   the episode's return
  "l":    <ndarray, int kind, shape [m]>,     its length in environment steps
  "s":    <ndarray "|b1" shape [m]>,          whether it succeeded
  "mask": <ndarray "|b1" shape [m]> }         whether this entry is a finished episode

The training algorithms (ppo, fpo, dppo) average the episodes recorded since the last learn step into rollout/reward, rollout/length and rollout/success, and start again after it. evaluation also logs each episode as it finishes. An entry whose mask is false is skipped, and a missing mask counts as true. Nothing else in info is read.

Leaving episode out is valid and changes nothing about training: the same transitions are stored and learned from. Only those metrics read 0. That is how it was found: a C++ client written from this document trained Pendulum normally while its logged return stayed at 0, until it sent episode (plugrl-server E44).

plugrl-env-client sends more than this, and the servers read none of the rest:

  • episode is full-length. Every environment in the batch has an entry; mask is true only where the episode ended at this step, and the other entries are zeros.
  • episode has a fifth key, mean_success_rate: a float over the client's last 100 episodes.
  • A top-level is_step_success of shape [m] rides on every feedback.

When the server looks for m in info, it also looks one level into a nested map. So {"episode": {...}} on its own is unbatched correctly, even though none of its top-level values is an ndarray.

Terminal observations

The client MUST send the observation that the terminal step returned, not the first observation of the next episode. This corresponds to Gymnasium's AutoresetMode.NEXT_STEP: the step reporting done returns the terminal observation, and the reset happens on the following call.

An environment that resets inside its own step() violates this silently — nothing raises, and the terminal transition simply carries the wrong observation into the learner's buffer. plugrl-env-client pins the declared mode of every environment it ships and demonstrates the corruption executably in its tests/test_terminal_observation.py.


6. Batching on the server

Not part of the wire format, but it determines what a client sees.

The server does not answer an infer immediately. It queues the request and a scheduler drains the queue into one batch, running the policy once for all of them. A request therefore waits for two things: the scheduler's polling interval (SCHEDULER_SLEEP_INTERVAL, 0.1 ms) and the batch-readiness condition — by default one queued request per connected client, or mini_infer_batch_size environments when that is set.

Two consequences for a client:

  • Latency is not independent of other clients. A client alone on a server is answered as fast as the scheduler polls. A client sharing one with slower peers waits for them, up to INFER_READY_TIMEOUT = 5 s, after which the server logs a warning and keeps waiting.
  • A client must not assume its request is the whole batch. Actions come back sliced out of a larger tensor; the slice boundaries are the request's own n.

7. Closing

7.1 Normal shutdown

When the algorithm signals that training is finished, the server closes with close code 1001 (going away) and reason:

plugrl-server-stop

A client seeing this reason should exit rather than reconnect, unless it was explicitly configured to wait for the next run.

7.2 Protocol error — resync

When a message fails validation — wrong message_type, missing required key — the server closes with 1001 (going away) and reason:

plugrl-server-resync

This means "your stream and mine are out of step". A client should reconnect, read the fresh metadata, and resume from a new infer. Any feedback it was holding is stale and MUST be dropped: the server's per-environment state went with the closed connection.

7.3 Feedback timeout

If 60 s pass after an action without a feedback, the server closes with 1001 and reason Feedback timeout. Environments slower than that must be split across more connections, or the constant raised on both sides.

7.4 Internal error

An unhandled server-side exception closes with 1011 (internal error) and reason Internal server error.. Unlike openpi (section 9), PlugRL does not send the traceback as a text frame first.

A client should nonetheless be prepared to receive a text frame where it expected binary, and treat it as a fatal server-side error rather than attempting to unpack it.

7.5 The client stopping

Everything above is the server closing. A client may also stop first - it has collected the episodes it was asked for, or its operator interrupted it - and that is not an error. The server keeps no state that outlives the connection, so a client that is done costs nothing beyond the feedback it had not yet sent. A client that intends to come back pays more than that; see section 7.6.

A client that is finished SHOULD send a WebSocket close frame with status 1000 (normal closure) before dropping the socket, as RFC 6455 section 5.5.1 asks. Nothing breaks without one - the server sees the connection end either way - but the difference is visible in its logs, as ConnectionClosedError: no close frame received or sent rather than a clean ConnectionClosedOK, and an operator reading those logs should not have to wonder whether a client crashed.

examples/conformance_server.py reports a missing close frame as a note, not a violation, which is the level this rule deserves.

Historical note. Until 2026-09-11 the C++ reference client did not send one. It went unnoticed because in every test until then the server ran out of steps first and closed the connection itself; E7, where the client finishes first, made it visible on all 45 runs.

7.6 Reconnecting

Section 7.2 says the server's per-environment state went with the closed connection. That is true of every close, not only a resync, and it is the one thing a reconnecting client has to reason about.

The server holds, per environment and for exactly one connection: the previous observation, the policy step state, and the terminated / truncated flags. A new connection begins with none of them. So:

  • a client that reconnects MUST drop any feedback it was holding;
  • it reads a fresh metadata and resumes from a fresh infer;
  • a feedback whose action arrived on an earlier connection describes a transition the server can no longer complete, because it has no previous observation to attach it to. Sending it produces a transition built from nothing, which is worse than the lost step it was trying to save.

A server SHOULD say something when it receives feedback for an environment it has no step state for. That condition has exactly one cause - the connection was replaced mid-run - and storing the transition silently puts a hole in the training data that nothing downstream can detect.

Gap - the reference client breaks this on one path. plugrl-env-client run with --reconnect-on-server-stop resends the feedback it was holding instead of dropping it, when the close it hit was the server's plugrl-server-stop. In websocket_env_client_agent.py, feedback()'s SERVER_STOP_REASON branch continues its retry loop, which opens a new connection, reads fresh metadata, and re-sends the same payload - the exact resend the historical note below says this section exists to prevent. Every other close path in that method returns and drops the transition, and tests/test_reconnect_drops_feedback.py covers three of them (a keepalive timeout, a plain 1000 close, a resync) but not this one. It is the only path in the reference client that violates the MUST above.

Where reconnects come from. Nothing in this protocol causes them and nothing in it can prevent them: a suspended laptop, a flaky link, an operator restarting the server. The rule above is written in terms of the reconnect rather than its cause, because the cost is the same either way, and because a client cannot tell the causes apart from where it sits.

Historical note. This section exists because of a run that dropped its connection after the machine it was on suspended for nearly two hours. The reconnect was handled, and the transition that crossed it was not: the env client resent the feedback it was holding, and the server completed it from an empty observation and stored it. The first diagnosis blamed WebSocket keepalive pings and a long learn step, which measurement then ruled out - plugrl-server runs learn off the event loop, and learns of 190 s produce no ping timeout. The cause was mundane. The hole in the protocol was not, and had nothing to do with it.


8. Conformance checklist

A client conforms to version 1 if it:

  • connects over WebSocket with compression disabled and no frame size cap;
  • reads one metadata message before sending anything, and requires no particular key in it;
  • emits binary frames containing msgpack maps with a message_type str;
  • encodes arrays with bin keys __ndarray__ / data / dtype / shape, C-contiguous, and parses the typestr rather than assuming a dtype;
  • sends exactly one feedback for every action received, and never two infer messages in a row;
  • tolerates its feedback env set differing from its infer env set;
  • reads action as [H, n, *da], time-major, and reads env_ids when present;
  • reports chunk-summed reward, and the terminal observation on the step that reports done;
  • handles close reasons plugrl-server-stop and plugrl-server-resync differently;
  • drops any held feedback when a connection closes, and never sends feedback for an action that arrived on an earlier connection;
  • treats a text frame as a fatal error;
  • uses a feature of section 10 only when the server lists it, and, for reuse-feedback-obs, reuses only an observation the server holds.

plugrl-conformance (examples/conformance_server.py is a shim for it) has two modes.

By default it watches one connection. It checks what that connection shows: framing, alternation, env indices, observation shape, and the feedback payload's keys, dtypes and lengths. It reports what it accepts but cannot require as a note rather than a failure. Both reference clients pass it with one note: they send text as a msgpack string array rather than a <U array, which is the section 3.4 Gap. An unexercised clause leaves no trace in its report, and from one connection to an unknown environment it cannot tell what the client did with an action.

With --probe the client runs the probe environment of section 8.1. From each feedback the checker then works out how many steps of the chunk the client ran and which action it applied last, and it drives the connection instead of watching it. In that mode it checks every clause of the list above except these:

  • reading env_ids: section 4.3 requires them to equal the infer's env_indices, so a client that ignores them behaves identically;
  • tolerating a feedback env set that differs from the infer set: the client chooses its sets, and the server cannot make them differ;
  • requiring no particular metadata key: the checker sends all of them;
  • dropping held feedback after a close other than a resync.

--scenario selects how the connection is driven:

  • basic: the checker waits 0.3 s before speaking, so a client that sends first is caught. Its metadata carries an unknown 2 MiB key, so a client that kept its library's 1 MiB frame cap cannot receive it, and one that rejects unknown keys fails. It checks that the handshake did not offer permessage-deflate. It answers every infer with float64, time-major actions whose values encode their position. Once the exchanges are done it closes with plugrl-server-stop, after which the client must exit 0 and not reconnect.
  • resync: as basic, but halfway through it closes with plugrl-server-resync right after sending an action, so the client is holding a feedback. The client must reconnect, and its first message on the new connection must be an infer.
  • text: it answers the first infer with a text frame, which the client must treat as fatal: it closes the connection and sends nothing more.
  • all: each in turn, starting the client once per scenario.

Both reference clients pass all three with the same note: examples/raw_client.py --probe and plugrl_client --probe, the C++ one, which CI checks on every change. The Python client's --bug option breaks one clause at a time, and tests/test_conformance_probe.py checks that the checker names each one.

8.1 The probe environment

A client that wants its handling of actions checked, and not only its messages, runs this environment for plugrl-conformance --probe. It needs no simulator, and is a few lines in any language.

  • The client runs n probe envs, with indices 0 to n - 1 as its env_indices. Indices stay below 1000.

  • Env i's episode lasts L_i = 3 + 2 * (i mod 3) steps: 3, 5, 7, 3, 5, ... So within one chunk, different envs end their episodes at different steps.

  • Its observation has no images and two states:

    • t: float64 [n, 1], the number of steps taken in the current episode, 0 after a reset;
    • a: float64 [n, d], the action the env applied on its last step, where d is the action width. Its value after a reset is not checked, and before the client knows the width it may have d = 1.

    text may be anything.

  • A step applies the action, adds 1 to t, stores the action in a, and pays a reward of 1. It reports terminated when t reaches L_i. The probe env never truncates.

  • A reset sets t back to 0.

The checker sends action values that encode their position: element [k, row, d] for env j is 10000 k + 10 j + d. A feedback therefore says exactly what the client did with the chunk. t minus the t of the infer it answered is the number of steps it ran, which must be between 1 and H. With a reward of 1 per step, the reward must equal that number, which is the chunk-sum rule. a must be the chunk's action at the last step run, for that env, which checks time-major order, the step order, and the dtype. A terminated env must report t = L_i, its own terminal observation, and its next infer must report t = 0.

8.2 Checking a server

The checklist above is for clients. A training server conforms if it:

  • sends one metadata message, before anything else, on every connection;
  • answers each infer with an action that is time-major, [H, n, *da], with env_ids equal to the infer's env_indices, in order, and that agrees with the action_horizon and action_dim it declared;
  • accepts what a client may send: a feedback env set unlike the infer's, an n that varies, and frames larger than 1 MiB;
  • closes a connection with 1001 and plugrl-server-resync on a malformed message, two infers in a row, or an info it cannot split per environment, and stays up for its other clients;
  • keeps env indices connection-scoped;
  • ends a run with 1001 and plugrl-server-stop;
  • if it lists reuse-feedback-obs, answers an infer that reuses observations, and closes with 1001 and plugrl-server-resync on one that reuses an observation it does not hold.

plugrl-conformance-server checks each of these from the client's side of the wire:

plugrl-conformance-server --port 8000 --state-dim 3 --until-stop

It sends a states-only observation, states["obs"] of --state-dim values (--state-key and --image-key change that), so it can drive any server whose policy reads one. examples/reference_server.py is a server written against this document that trains nothing and passes every check. Its --bug option breaks one clause at a time, and tests/test_server_conformance.py checks that the grader names each one. plugrl-server passes too. Its version from before plugrl-server#108 fails two checks: an info it could not split closed that connection with 1011, and the server then refused new connections.

What the server does with a frame is not visible from the wire, so section 7.6's SHOULD - saying something about feedback for an environment it holds no step state for - is not checked.


9. Relationship to openpi

PlugRL's serialization is openpi's. msgpack_numpy.py is taken from openpi (Apache-2.0, attributed in the file header and in NOTICE), reformatted only. The array encoding is byte-identical, deliberately, so that the hard part of writing a client is shared between the two ecosystems.

The message layer differs:

openpi PlugRL
First server message the metadata map, bare {"message_type": "metadata", "data": ...}
Request the observation map, bare envelope + env_indices + step_ids
Response the action map, bare, plus server_timing envelope + env_ids + action
Return channel none feedback: reward, done flags, next observation
Batching one observation per request ragged multi-environment batch
Errors traceback sent as a text frame, then close 1011 close 1011, no traceback frame
Health check GET /healthz none

PlugRL is openpi's serving protocol plus a feedback return channel. That is the whole difference, and it is the difference between serving a policy and training one. An openpi client cannot drive a PlugRL server unmodified, and vice versa, but the array codec — the part that takes real work in C++ or Rust — carries across unchanged.


10. Versioning

This is version 1. The server publishes it as protocol_version in the metadata message (section 5.1). A client MUST NOT require the key, and the server does not adapt to a client's version.

What is negotiated is features. A feature is an addition to version 1 that a client may use and a server may offer, and nothing else changes for either side when it is absent. The server lists the ones it implements in metadata's features. A client uses a feature only if it is listed there, and a server that lists one still accepts every client that does not use it. So a new client works with an old server, and an old client with a new one. A change that cannot be made this way is breaking, and moves protocol_version to 2.

Correction, 2026-09-11. This paragraph used to say version 1 "has no version field on the wire" and attributed to section 5.1 a remark that the metadata message was the obvious place for one. Both halves went stale on 2026-09-10, when the metadata message gained contents: protocol_version has been in it on every connection since, in plugrl-server and in examples/conformance_server.py alike, and section 5.1 lists it as a key the server sends rather than as a suggestion. What survives is the absence of negotiation, which is what the sentence was reaching for.

Changing any of the four message_type strings, the two close reasons, the ndarray key names, the action array's axis order, or the meaning of rewards is a breaking change.

The first three are pinned in this repository, and a change to them fails a test rather than a deployment: tests/test_wire_format.py holds the four message_type strings and the two close reasons, tests/test_spec_conformance.py the four ndarray key names. The last two are not, and cannot be - this is the codec repository, and the axis order and the chunk-sum rule are properties of how the two sides behave, not of what the packer emits. Transposing the action chunk or redefining rewards as the last step's reward passes every codec test here. They are checked instead by plugrl-conformance --probe (section 8.1), against any client.

10.1 reuse-feedback-obs

Without it, every observation crosses the link twice. The client sends it in the feedback that ends a chunk, as the step's next observation, and then sends the same observation again in the next infer, as the state the new chunk starts from. Only after a reset do the two differ. That second copy is the factor of two in plugrl-server's E43 cost model: an exchange costs a fixed latency plus twice the observation's bytes over the link.

With the feature, the infer may say "the observation you already have":

{ "message_type": "infer",
  "data":        <observation, batched to the rows that are not reused>,
  "env_indices": <ndarray "<i8" shape [n]>,
  "step_ids":    <ndarray "<i8" shape [n]>,
  "reuse":       <ndarray "|b1" shape [n]> }
  • reuse[i] true means env env_indices[i] sends no observation. Its observation is the obs row of the last feedback that named it on this connection.
  • data is batched to the rows where reuse is false, in their env_indices order. With every row reused, its arrays have a leading dimension of 0.
  • A client MUST send reuse false for an env whose last feedback on this connection reported terminated or truncated, since that feedback carried the terminal observation and not the reset one. It MUST also send it false for an env with no feedback on this connection yet, which includes every env just after a reconnect.
  • A server that lists the feature MUST close a connection whose infer reuses an observation it does not hold - an env never fed back on this connection, or one whose last feedback ended its episode - with 1001 and plugrl-server-resync.
  • An infer without reuse sends every observation, as in version 1.

Both checkers know the feature. plugrl-conformance --features reuse-feedback-obs offers it to a client and checks how the client uses it. plugrl-conformance-server exercises it when the server lists it.