Skip to content
@PlugRL

PlugRL

PlugRL

Reinforcement learning training and the environments it learns from, split into two processes and joined by a written protocol: WebSocket and msgpack, with a feedback return channel.

Two boxes. On the left, plugrl-env-client, which holds no policy and runs Gymnasium, MuJoCo or LIBERO. On the right, plugrl-server, which holds the policy and the algorithm. Observation goes from client to server, action from server to client, and feedback - reward and termination - from client to server, drawn thicker and green.

The third arrow is the one that matters. Serving an inference model needs the first two; learning from what happened needs the third, and the protocol specifies it rather than leaving it to a convention.

What runs on it

Sixteen cells, four policy-algorithm pairs on HalfCheetah, Hopper, Walker2d and robomimic square, each with a frame from its trained policy, a training curve and a status. Every pair learns every task. On square each starts from a pretrained policy, and fpo-policy with FPO passes the bar on two of three seeds.

Every combination of the two MLP policies and the two algorithms on four tasks, and the baseline they are measured against, a Gaussian MLP with PPO. All sixteen learn. On the project page each cell plays its clip and shows the two commands that trained it. None of the servers that trained these has MuJoCo, robosuite or gymnasium installed; the env clients carry them, in two separate environments.

What the split buys, and what it costs

The training server is 6.5G and wants a GPU. The environment side needs neither, and need not be Python: it fits on a different class of machine from the trainer. The boundary between them costs a fixed latency plus the observation's bytes over the link, small on a fast link and measurable on a slow one.

Question Answer
E44 Does an env client have to be this codebase, or Python? No - a C++ program with no third-party libraries trains a policy on its own Pendulum as well as the Python env client does (E2 first spoke the protocol from C++)
E12 Does a rollout machine need CUDA? No - LIBERO's env client goes from 7.8G to 3.4G, with no nvidia wheels
E13 Or a GPU to render on? No, at 1.91x the wall clock - ten clients rendering on the CPU, 30 of 30 episodes successful
E43 Does training still work with the env clients on another physical machine? Yes - the quickstart pair learns with its env clients on a Windows laptop over campus Wi-Fi
E43 What does crossing cost? About 3 ms plus twice the observation's bytes over the link per exchange: 21 ms for 184 KiB at 18 MB/s
E10 Is that cheap beside a VLA forward pass? On a fast link. Over that Wi-Fi a 184 KiB observation is 21% of pi0.5's 100 ms forward, not E10's 1.3-3.6%

Every experiment directory carries its data and a FINDINGS.md that states what the result does not support.

A real VLA through it

A full-size pi0.5 runs end to end through the boundary on LIBERO. The unmodified checkpoint scored 99 of 100 on libero_spatial and 185 of 200 on libero_10, against openpi's published 98.8 and 92.4, and the server's record of episodes and steps reconciles exactly with the clients' (E11). Fine-tuning it with reinforcement learning through PlugRL has not made it better yet; that record is on its own page.

Repositories

plugrl-server Training side: policy, algorithm, checkpoints, and the experiments
plugrl-env-client Environment side: steps envs, asks for actions, returns feedback
plugrl-protocol The specification, its checkable clauses as tests, and two reference env clients
plugrl.github.io Documentation, in English and 中文

Start at the documentation - the quickstart trains FPO on HalfCheetah with no GPU and nothing to download.

Who

PlugRL is built by Chenhao Lu, Zuo Gou and Zilin Kang.

Popular repositories Loading

  1. plugrl-server plugrl-server Public

    Training server for PlugRL - policy and algorithm behind a WebSocket boundary with a feedback return channel. A full-size pi0.5 ran end to end through it on LIBERO: inference, feedback and FPO trai…

    Python

  2. plugrl.github.io plugrl.github.io Public

    Documentation for PlugRL, in English and 中文 - quickstarts, the protocol specification, and what the experiments found.

  3. plugrl-protocol plugrl-protocol Public

    The written boundary between RL training and its environments: WebSocket + msgpack, with a feedback return channel. SPEC.md is normative and its checkable clauses are tests; two reference env clien…

    Python

  4. plugrl-env-client plugrl-env-client Public

    Environment side of PlugRL - runs Gymnasium, MuJoCo and LIBERO environments and asks a training server for actions over WebSocket. It holds no policy and no training stack.

    Python

  5. .github .github Public

    Organization profile for PlugRL.

Repositories

Showing 5 of 5 repositories

Top languages

Loading…

Most used topics

Loading…