Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
name: CI

# deploy-pages.yml only runs on a push to main, so until now a pull request
# got no validation at all - a broken build was only visible after merging.
on:
push:
branches: [main]
pull_request:
workflow_dispatch:

concurrency:
group: ci-${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v5
with:
enable-cache: true
- name: Sync
run: uv sync --frozen
# --strict turns warnings into failures, so a broken link or a page
# missing from the nav fails here rather than shipping.
- name: Build
run: uv run mkdocs build --strict
2 changes: 1 addition & 1 deletion docs/env/custom_env.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,4 +75,4 @@ plugrl-run-env-client custom-v1 --num-episodes 1
## Next steps

- [Environments](index.md)
- [Remote viewer](../user_guide/get_started.md)
- [Get Started](../user_guide/get_started.md)
1 change: 0 additions & 1 deletion docs/env/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,6 @@ env = gym.make(env_id, config=config_dataclass, max_episode_steps=max_episode_st

- `--num-workers`: run multiple env client processes
- `--server-host`, `--server-port`: server address
- `--use-remote-viewer`: stream observations to the viewer
- `--use-real-time`, `--fps`: fixed FPS for debugging

## Troubleshooting
Expand Down
1 change: 0 additions & 1 deletion docs/env/index.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,6 @@ env = gym.make(env_id, config=config_dataclass, max_episode_steps=max_episode_st

- `--num-workers`:多进程并行跑环境
- `--server-host`、`--server-port`:server 地址
- `--use-remote-viewer`:推送观测到 viewer
- `--use-real-time`、`--fps`:固定 FPS 运行

## 常见问题
Expand Down
59 changes: 52 additions & 7 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,35 +6,80 @@ PlugRL is an RL infrastructure for distributed experiments with a clean split be

## Quickstart

Run a connectivity smoke test.
Two processes: a training server that holds the policy, and an env client
that runs environments and asks it for actions. This pair actually learns -
FPO on HalfCheetah-v5, CPU only, no GPU and no assets to download.

```bash
# Terminal 1 - the training server
plugrl-run-server fpo-policy default fpo default \
--port 8000 --policy.device cpu \
--algo.global-steps 500000 --algo.buffer-size 4096

# Terminal 2 - the environment
plugrl-run-env-client mujoco-v1 \
--server-host 127.0.0.1 --server-port 8000 \
--num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
```

Episode return climbs out of the -300s within a few minutes. `HalfCheetah-v5`
has a 17-dimensional observation and a 6-dimensional action, which are
exactly `fpo-policy`'s defaults, so nothing needs configuring. The
environment needs `plugrl-env-client[mujoco]`.

!!! warning "`--algo.buffer-size` is not decoration"

FPO learns when its rollout buffer fills, or when the run reaches its
last step. At the default `buffer_size=983040`, a run shorter than about
a million steps therefore learns **exactly once, at the very end** -
which gives you a single point instead of a curve.

### Just checking connectivity?

```bash
plugrl-run-server dummy-policy default dummy default
plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
```

The dummy algorithm's `learn` is a sleep - it moves no weights. Use it to
confirm the two sides talk to each other, not to train anything.

## Verify

- Server prints a WebSocket listening address.
- Env client prints server metadata and starts stepping episodes.
- The server prints a WebSocket listening address.
- The env client prints the server's metadata - policy name, action shape -
and starts stepping episodes.
- With `fpo`, the server prints a metrics table whose `rollout/reward` rises.

## Components

- `plugrl-server`: training server, runs algorithm, policy, checkpoints, tracking
- `plugrl-env-client`: environment runner, collects rollouts
- `plugrl-protocol`: transport, message types, and serialization (WebSocket + msgpack)
- `plugrl-monitor`: optional remote viewer for observations

The boundary between the first two is [the protocol](protocol/index.md), and
it is specified rather than implied: an env client does not have to be
Python, or be this codebase.

## Common options

- Env client connects to server via `--server-host` and `--server-port`.
- Stream observations with `--use-remote-viewer`, `--viewer-host`, `--viewer-port`.
- Use `plugrl-run-server-ray` for Ray-based distributed launch.
- The env client connects to the server via `--server-host` and `--server-port`.
- `--num-procs` runs several env client processes against one server.

!!! note "On `plugrl-run-server-ray`"

There is a Ray-based launcher, but it is **not a supported path today**.
It requires the `dppo` extra, builds its worker list from the *local*
GPU count so a multi-node cluster still only sees the head node, and its
server speaks an older dialect of the protocol than the WebSocket one -
see [SPEC.md section 5.3](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md).
Use `plugrl-run-server` unless you are working on the Ray path itself.

## Next steps

- [User Guide](user_guide/index.md)
- [Get Started](user_guide/get_started.md)
- [Protocol](protocol/index.md)
- [Algorithms](algorithm/index.md)
- [Environments](env/index.md)
- [Policies](policy/index.md)
Expand Down
49 changes: 44 additions & 5 deletions docs/index.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,35 +6,74 @@ PlugRL 是一套面向分布式强化学习实验的基础设施。训练端与

## 快速开始

先跑通一次联通性 smoke test。
两个进程:训练端持有策略,环境端跑环境并向它请求动作。下面这一对**真的会学** ——
FPO + HalfCheetah-v5,纯 CPU,不需要 GPU,也不需要下载任何资源文件。

```bash
# 终端 1 —— 训练端
plugrl-run-server fpo-policy default fpo default \
--port 8000 --policy.device cpu \
--algo.global-steps 500000 --algo.buffer-size 4096

# 终端 2 —— 环境端
plugrl-run-env-client mujoco-v1 \
--server-host 127.0.0.1 --server-port 8000 \
--num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
```

几分钟内 episode 回报就会从 -300 附近爬上来。`HalfCheetah-v5` 的观测是 17 维、
动作是 6 维,**正好是 `fpo-policy` 的默认值**,所以不需要任何配置。
环境端需要 `plugrl-env-client[mujoco]`。

!!! warning "`--algo.buffer-size` 不是装饰"

FPO 在 rollout buffer 填满时、或运行到最后一步时才学习。按默认的
`buffer_size=983040`,任何少于约一百万步的运行**只会在最后学一次** ——
你得到的是一个点,不是一条曲线。

### 只想确认能连通?

```bash
plugrl-run-server dummy-policy default dummy default
plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
```

dummy 算法的 `learn` 是一个 sleep,不会移动任何权重。它用来确认两端能对话,
不是用来训练的。

## 验证

- server 打印 WebSocket 监听地址
- env client 打印 server 元信息并开始跑 episode
- env client 打印 server 元信息(策略名、动作形状)并开始跑 episode
- 用 `fpo` 时,server 的指标表里 `rollout/reward` 会上升

## 组件

- `plugrl-server`:训练端,负责算法、策略、checkpoint、指标追踪
- `plugrl-env-client`:环境端,负责创建环境并采集 rollout
- `plugrl-protocol`:协议与序列化层,WebSocket 与 msgpack
- `plugrl-monitor`:可选 viewer,用于查看观测

前两者之间的边界就是[通信协议](protocol/index.zh.md),而且它是**被写下来的**而非
默认的:环境端不必是 Python,也不必是这个代码库。

## 常用参数

- env client 通过 `--server-host` 与 `--server-port` 连接 server
- 观测串流使用 `--use-remote-viewer`、`--viewer-host`、`--viewer-port`
- 分布式启动使用 `plugrl-run-server-ray`
- `--num-procs` 可以起多个环境端进程连同一个 server

!!! note "关于 `plugrl-run-server-ray`"

确实有一个基于 Ray 的启动器,但**目前不是受支持的路径**。它需要 `dppo`
extra;它用**本机**的 GPU 数构建 worker 列表,所以即使 Ray 连上多节点集群
也只看得见头节点;而且它的服务端说的是比 WebSocket 版更旧的协议方言,
见 [SPEC.md 第 5.3 节](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)。
除非你就是在改 Ray 这条路径,否则请用 `plugrl-run-server`。

## 下一步

- [用户指南](user_guide/index.zh.md)
- [快速开始](user_guide/get_started.zh.md)
- [通信协议](protocol/index.zh.md)
- [算法](algorithm/index.zh.md)
- [环境](env/index.zh.md)
- [策略](policy/index.zh.md)
Expand Down
95 changes: 95 additions & 0 deletions docs/protocol/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# The wire protocol

PlugRL splits a training run across two processes. A **training server**
holds the policy and the learning algorithm. An **env client** runs
environments, asks for actions, and reports what happened. They talk over
WebSocket, with msgpack on the wire.

That boundary is the reason an environment stack and a training stack never
have to share a Python environment — and the reason an env client does not
have to be Python at all. A ROS node can be one. So can a robot's onboard
C++ controller.

!!! info "The specification lives in `plugrl-protocol`"

**[SPEC.md](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)**
is the normative document. It is kept next to the code it describes so
the two cannot drift apart, and this page only orients you.

It also names its own known defects, in boxes marked **Gap**. Those are
the honest part; read them before building on anything.

## The exchange

```
client server
|------------ WebSocket handshake ---------->|
|<---------------- metadata -----------------| the server speaks first
|------------------ infer ------------------>| observations
|<----------------- action ------------------| an action chunk
|----------------- feedback ---------------->| reward, done, next obs
```

Four message types — `metadata`, `infer`, `action`, `feedback` — carried as
msgpack maps with a `message_type` field. Arrays travel as

```
{b"__ndarray__": true, b"data": <bin>, b"dtype": "<f4", b"shape": [4, 1, 7]}
```

`dtype` is a numpy typestr: a byte-order character, a kind character, and an
item size. Parsing it takes about ten lines in any language.

## Three rules a first implementation usually gets wrong

**Messages strictly alternate.** `infer`, `action`, `feedback`, `infer`, and
so on. The server's connection handler is straight-line code with no
dispatcher, so a client that sends two `infer` messages in a row has the
second one parsed as a `feedback` and is disconnected.

**The environment sets in one cycle need not match.** `infer` carries the
environments whose action chunk has run out; `feedback` carries the ones
whose chunk finished on this step. The first time an environment terminates
early, those stop being the same set — permanently. The pairing between an
`action` and the `feedback` after it is flow control, not association; the
server routes feedback by environment index.

**The reward is the sum over the chunk.** Not the last step's. A client that
reports the final step's reward trains a different MDP, and nothing fails.

## Checking an implementation

`plugrl-protocol` ships a server that grades a client against the
specification clause by clause and exits non-zero on a violation:

```bash
python examples/conformance_server.py --port 8000 --steps 20 &
./plugrl_client 127.0.0.1 8000 20
```

Its report has two severities. A **violation** is something the real server
would reject or mishandle. A **note** is something it accepts that differs
from what the Python client does — a portability risk, not a breach.

Two reference clients pass it. `raw_client.py` is 275 lines of Python using
only `msgpack` and `websockets` — no numpy, nothing from PlugRL.
`plugrl_client.cpp` is C++17 with **no third-party libraries at all**: SHA-1,
base64, WebSocket framing and the msgpack subset the protocol needs are all
in the one file, because that is the situation an embedded controller is
actually in.

Both run in CI on every change, so the claim on this page stays checked
rather than remembered.

## Relationship to openpi

PlugRL's serialization is
[openpi](https://github.com/Physical-Intelligence/openpi)'s. The
`msgpack_numpy` codec is taken from it under Apache-2.0 and reformatted
only, so **the array encoding is byte-identical** — the part of a client
that takes real work in C++ or Rust carries across between the two.

The message layer differs. openpi sends a bare observation and gets a bare
action back, which is what serving a policy needs. PlugRL wraps both in an
envelope and adds the `feedback` return channel, which is what training one
needs.
79 changes: 79 additions & 0 deletions docs/protocol/index.zh.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# 通信协议

PlugRL 把一次训练拆成两个进程。**训练服务端**持有策略与学习算法,**环境客户端**
运行环境、请求动作、汇报结果。两者通过 WebSocket 通信,线上格式是 msgpack。

这条边界正是训练栈与环境栈不必共处同一个 Python 环境的原因,也是环境客户端
根本不必是 Python 的原因 —— 一个 ROS 节点可以是,机器人上的 C++ 控制器也可以是。

!!! info "规范正本在 `plugrl-protocol`"

**[SPEC.md](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)**
是规范性文件。它与所描述的代码放在一起,以免两者漂移;本页只做导览。

规范里用 **Gap** 标注了协议自己的已知缺陷。那部分才是诚实的部分,
在此之上开发前请先读它们。

## 一次交换

```
客户端 服务端
|------------ WebSocket 握手 -------------->|
|<---------------- metadata -----------------| 服务端先说话
|------------------ infer ------------------>| 观测
|<----------------- action ------------------| 一段动作块
|----------------- feedback ---------------->| 奖励、终止标志、下一观测
```

四种消息类型 —— `metadata`、`infer`、`action`、`feedback` —— 都是带
`message_type` 字段的 msgpack map。数组的形式是

```
{b"__ndarray__": true, b"data": <bin>, b"dtype": "<f4", b"shape": [4, 1, 7]}
```

`dtype` 是 numpy 的 typestr:一个字节序字符、一个类型字符、一个元素字节数。
用任何语言解析它大约十行代码。

## 三条最容易实现错的规则

**消息严格交替。** `infer`、`action`、`feedback`、`infer`……服务端的连接处理
是一段没有分发器的顺序代码,所以连发两个 `infer` 的客户端会让第二个被当成
`feedback` 解析,然后被断开。

**一个周期里三条消息的环境集合不必相同。** `infer` 携带的是动作块刚用完的
环境,`feedback` 携带的是本步结束时动作块用完的环境。只要有一个环境提前终止,
这两个集合就**永久**不再相等。`action` 与其后 `feedback` 的配对只是流控,
不是语义关联 —— 服务端按环境编号查表路由。

**奖励是整个动作块上的求和**,不是最后一步的奖励。只汇报最后一步的客户端会在
一个不同的 MDP 上训练,而且不会有任何东西报错。

## 检验一个实现

`plugrl-protocol` 附带一个服务端,它按规范逐条给客户端打分,有违规就以非零码退出:

```bash
python examples/conformance_server.py --port 8000 --steps 20 &
./plugrl_client 127.0.0.1 8000 20
```

报告分两个等级。**violation** 是真服务端会拒绝或处理错的问题;**note** 是真服务端
接受、但与 Python 客户端做法不同的地方 —— 是可移植性风险,不是违约。

两个参考客户端都能通过。`raw_client.py` 是 275 行 Python,只用 `msgpack` 和
`websockets`,不用 numpy,也不用 PlugRL 的任何东西。`plugrl_client.cpp` 是 C++17,
**完全不依赖第三方库**:SHA-1、base64、WebSocket 分帧,以及协议需要的那部分
msgpack,全都写在同一个文件里 —— 因为嵌入式控制器面对的就是这种处境。

两者都在每次改动的 CI 里跑,所以本页的说法是被持续检验的,而不是被记住的。

## 与 openpi 的关系

PlugRL 的序列化就是
[openpi](https://github.com/Physical-Intelligence/openpi) 的。`msgpack_numpy`
取自 openpi(Apache-2.0),仅重新排版,因此**数组编码逐字节相同** —— 用 C++ 或
Rust 写客户端时真正费力的那部分,可以在两个生态之间通用。

不同之处在消息层。openpi 发一个裸观测、收一个裸动作,这是**服务**一个策略所需的;
PlugRL 给两者加了信封,并加上 `feedback` 回传通道,这是**训练**一个策略所需的。
Loading
Loading