From aad70179c95b0717fe25c6122e36da1680cfd744 Mon Sep 17 00:00:00 2001 From: Gotham-Zolio <18781106300@163.com> Date: Wed, 9 Sep 2026 15:44:46 -0400 Subject: [PATCH 1/6] docs: remove references to software that does not exist The site documented a remote viewer across seven pages - a uvicorn app named vlarl_viewer, a plugrl-monitor component, and --use-remote-viewer/--viewer-host /--viewer-port flags. None of it exists: there is no viewer repository, and the env client CLI has no such flags. Readers following these instructions could only fail. Also renames the remaining "worker" wording to "env client" to match the package that ships today. Co-Authored-By: Claude Opus 5 (1M context) --- docs/env/custom_env.md | 2 +- docs/env/index.md | 1 - docs/env/index.zh.md | 1 - docs/index.md | 2 -- docs/index.zh.md | 2 -- docs/user_guide/get_started.md | 19 ------------------- docs/user_guide/get_started.zh.md | 21 --------------------- docs/user_guide/index.md | 5 +---- docs/user_guide/index.zh.md | 5 +---- 9 files changed, 3 insertions(+), 55 deletions(-) diff --git a/docs/env/custom_env.md b/docs/env/custom_env.md index 4d48ded..f909e56 100644 --- a/docs/env/custom_env.md +++ b/docs/env/custom_env.md @@ -75,4 +75,4 @@ plugrl-run-env-client custom-v1 --num-episodes 1 ## Next steps - [Environments](index.md) -- [Remote viewer](../user_guide/get_started.md) \ No newline at end of file +- [Get Started](../user_guide/get_started.md) \ No newline at end of file diff --git a/docs/env/index.md b/docs/env/index.md index ae8649c..87d3ce1 100644 --- a/docs/env/index.md +++ b/docs/env/index.md @@ -43,7 +43,6 @@ env = gym.make(env_id, config=config_dataclass, max_episode_steps=max_episode_st - `--num-workers`: run multiple env client processes - `--server-host`, `--server-port`: server address -- `--use-remote-viewer`: stream observations to the viewer - `--use-real-time`, `--fps`: fixed FPS for debugging ## Troubleshooting diff --git a/docs/env/index.zh.md b/docs/env/index.zh.md index edee597..f6078c0 100644 --- a/docs/env/index.zh.md +++ b/docs/env/index.zh.md @@ -45,7 +45,6 @@ env = gym.make(env_id, config=config_dataclass, max_episode_steps=max_episode_st - `--num-workers`:多进程并行跑环境 - `--server-host`、`--server-port`:server 地址 -- `--use-remote-viewer`:推送观测到 viewer - `--use-real-time`、`--fps`:固定 FPS 运行 ## 常见问题 diff --git a/docs/index.md b/docs/index.md index 8f534de..49925c0 100644 --- a/docs/index.md +++ b/docs/index.md @@ -23,12 +23,10 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server - `plugrl-server`: training server, runs algorithm, policy, checkpoints, tracking - `plugrl-env-client`: environment runner, collects rollouts - `plugrl-protocol`: transport, message types, and serialization (WebSocket + msgpack) -- `plugrl-monitor`: optional remote viewer for observations ## Common options - Env client connects to server via `--server-host` and `--server-port`. -- Stream observations with `--use-remote-viewer`, `--viewer-host`, `--viewer-port`. - Use `plugrl-run-server-ray` for Ray-based distributed launch. ## Next steps diff --git a/docs/index.zh.md b/docs/index.zh.md index dbbd89f..1d4edd2 100644 --- a/docs/index.zh.md +++ b/docs/index.zh.md @@ -23,12 +23,10 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server - `plugrl-server`:训练端,负责算法、策略、checkpoint、指标追踪 - `plugrl-env-client`:环境端,负责创建环境并采集 rollout - `plugrl-protocol`:协议与序列化层,WebSocket 与 msgpack -- `plugrl-monitor`:可选 viewer,用于查看观测 ## 常用参数 - env client 通过 `--server-host` 与 `--server-port` 连接 server -- 观测串流使用 `--use-remote-viewer`、`--viewer-host`、`--viewer-port` - 分布式启动使用 `plugrl-run-server-ray` ## 下一步 diff --git a/docs/user_guide/get_started.md b/docs/user_guide/get_started.md index a643e4d..4387918 100644 --- a/docs/user_guide/get_started.md +++ b/docs/user_guide/get_started.md @@ -23,24 +23,6 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server - Server prints a WebSocket listening address. - Env client prints server metadata and steps episodes. -## Remote viewer - -Terminal C starts the viewer. - -```bash -uvicorn vlarl_viewer.main:app --reload --host 0.0.0.0 --port 9000 -``` - -Open this URL. - -- `http://localhost:9000/` - -Run the env client with streaming enabled. - -```bash -plugrl-run-env-client dummy-v1 --use-remote-viewer --viewer-host 127.0.0.1 --viewer-port 9000 -``` - ## DPPO examples Single process server. @@ -64,7 +46,6 @@ plugrl-run-server-ray dppo-policy default dppo hopper --exp_name my_dppo_exp --n ## Troubleshooting - Env client keeps retrying: check server address and firewall. -- Viewer shows disconnected: check `--viewer-host` and `--viewer-port`, then confirm `--use-remote-viewer` is enabled. - `--resume` requires existing checkpoints. ## Next steps diff --git a/docs/user_guide/get_started.zh.md b/docs/user_guide/get_started.zh.md index a390cdd..fba7a79 100644 --- a/docs/user_guide/get_started.zh.md +++ b/docs/user_guide/get_started.zh.md @@ -23,26 +23,6 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server - server 打印 WebSocket 监听地址 - env client 打印 server 元信息并开始跑 episode -## 远程 Viewer - -终端 C 启动 viewer。 - -```bash -uvicorn vlarl_viewer.main:app --reload --host 0.0.0.0 --port 9000 -``` - -打开这个地址。 - -- `http://localhost:9000/` - -env client 启用推流。 - -```bash -plugrl-run-env-client dummy-v1 --use-remote-viewer --viewer-host 127.0.0.1 --viewer-port 9000 -``` - -> Note: 多进程 env client 时只有进程 0 建立 viewer 连接,其它进程复用该连接。 - ## DPPO 示例 单进程 server。 @@ -66,7 +46,6 @@ plugrl-run-server-ray dppo-policy default dppo hopper --exp_name my_dppo_exp --n ## 常见问题 - env client 一直重试:检查 server 是否已启动,host 与 port 是否一致,端口是否可达。 -- viewer 显示未连接:检查 `--viewer-host` 与 `--viewer-port`,并确认已启用 `--use-remote-viewer`。 - `--resume` 找不到 checkpoint:确认实验目录存在且包含 checkpoint。 ## 下一步 diff --git a/docs/user_guide/index.md b/docs/user_guide/index.md index 2bc9d2c..5f5d7a0 100644 --- a/docs/user_guide/index.md +++ b/docs/user_guide/index.md @@ -1,6 +1,6 @@ # User Guide -Run PlugRL end to end: start a server, start workers, and stream observations when debugging. +Run PlugRL end to end: start a server, then start one or more env clients. ## Quickstart @@ -18,20 +18,17 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server 1. Start a training server with `plugrl-run-server` or `plugrl-run-server-ray`. 2. Start one or more env clients with `plugrl-run-env-client `. -3. Stream observations to the viewer when you need to debug. ## Components - `plugrl-server`: batches inference across connected workers, runs learning and checkpointing - `plugrl-env-client`: creates Gymnasium envs, sends `infer`, receives `action`, sends `feedback` - `plugrl-protocol`: WebSocket transport, message types, and msgpack serialization -- `plugrl-monitor`: optional remote viewer ## Common options - Server default address is `0.0.0.0:8000`. - Env client connects via `--server-host` and `--server-port`. -- Viewer streaming uses `--use-remote-viewer`, `--viewer-host`, `--viewer-port`. ## Troubleshooting diff --git a/docs/user_guide/index.zh.md b/docs/user_guide/index.zh.md index 253a95e..1ebe4c4 100644 --- a/docs/user_guide/index.zh.md +++ b/docs/user_guide/index.zh.md @@ -1,6 +1,6 @@ # 用户指南 -端到端跑通 PlugRL:启动 server,启动 worker,调试时把观测推到 viewer。 +端到端跑通 PlugRL:启动 server,启动 env client。 ## 快速开始 @@ -18,20 +18,17 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server 1. 用 `plugrl-run-server` 或 `plugrl-run-server-ray` 启动训练端。 2. 用 `plugrl-run-env-client ` 启动一个或多个环境端。 -3. 需要看观测时开启 viewer 推流。 ## 组件 - `plugrl-server`:聚合推理请求,驱动学习与 checkpoint - `plugrl-env-client`:创建 Gymnasium 环境,发送 `infer`,接收 `action`,回传 `feedback` - `plugrl-protocol`:WebSocket 传输与 msgpack 序列化 -- `plugrl-monitor`:可选 viewer ## 常用参数 - server 默认地址为 `0.0.0.0:8000` - env client 通过 `--server-host` 与 `--server-port` 连接 -- viewer 推流使用 `--use-remote-viewer`、`--viewer-host`、`--viewer-port` ## 常见问题 From b26dcdfaafa103514e45d12c974d4633f3738654 Mon Sep 17 00:00:00 2001 From: Gotham-Zolio <18781106300@163.com> Date: Wed, 9 Sep 2026 15:47:26 -0400 Subject: [PATCH 2/6] fix: give the docs site syntax highlighting and working nav metadata mkdocs.yml declared no markdown_extensions at all, so every fenced block on a site made almost entirely of shell and Python rendered as bare
 -
no highlighting, and nothing for content.code.copy to attach a button to.

theme.features listed "navigation.instance", which mkdocs-material does not
have, so it was silently ignored. Correcting it to navigation.instant turns out
to break the build: mkdocs-static-i18n cannot keep the language switcher
contextual with instant loading on, and --strict fails. The typo was
accidentally load-bearing. Left out, with a comment saying why.

Also adds site_description, repo_url and edit_uri, so pages get an
"Edit this page" link.

Co-Authored-By: Claude Opus 5 (1M context) 
---
 mkdocs.yml | 24 +++++++++++++++++++++++-
 1 file changed, 23 insertions(+), 1 deletion(-)

diff --git a/mkdocs.yml b/mkdocs.yml
index 8c16cc8..027d5dd 100644
--- a/mkdocs.yml
+++ b/mkdocs.yml
@@ -1,5 +1,9 @@
 site_url: https://plugrl.github.io
 site_name: PlugRL
+site_description: An RL infrastructure that keeps the training stack and the environment stack in separate dependency worlds.
+repo_url: https://github.com/PlugRL/plugrl.github.io
+repo_name: PlugRL/plugrl.github.io
+edit_uri: edit/main/docs/
 plugins:
   - search
   - i18n:
@@ -33,10 +37,28 @@ theme:
     - navigation.tabs.sticky
     - navigation.expand
     - navigation.footer
-    - navigation.instance
+    # navigation.instant is deliberately absent: mkdocs-static-i18n cannot keep
+    # the language switcher contextual with it enabled, and --strict fails.
     - toc.follow
+    - content.code.copy
+    - content.action.edit
     - search.suggest
     - search.highlight
+
+# Without these, fenced blocks render as bare 
: no syntax
+# highlighting, and content.code.copy has nothing to attach a button to.
+markdown_extensions:
+  - admonition
+  - attr_list
+  - toc:
+      permalink: true
+  - pymdownx.highlight:
+      anchor_linenums: true
+      pygments_lang_class: true
+  - pymdownx.inlinehilite
+  - pymdownx.superfences
+  - pymdownx.details
+
 nav:
   - Home: index.md
 

From 73ebb4baa9d893b2261762520c6685f765f69a15 Mon Sep 17 00:00:00 2001
From: Gotham-Zolio <18781106300@163.com>
Date: Thu, 10 Sep 2026 12:29:11 -0400
Subject: [PATCH 3/6] ci: validate the docs build on pull requests

deploy-pages.yml only runs on a push to main, so a pull request got no
validation at all and a broken build was first visible after merging.
--strict makes a warning - a dead link, a page missing from the nav - fail
the check rather than ship.

Co-Authored-By: Claude Opus 5 (1M context) 
---
 .github/workflows/ci.yml | 28 ++++++++++++++++++++++++++++
 1 file changed, 28 insertions(+)
 create mode 100644 .github/workflows/ci.yml

diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
new file mode 100644
index 0000000..fee5e6a
--- /dev/null
+++ b/.github/workflows/ci.yml
@@ -0,0 +1,28 @@
+name: CI
+
+# deploy-pages.yml only runs on a push to main, so until now a pull request
+# got no validation at all - a broken build was only visible after merging.
+on:
+  push:
+    branches: [main]
+  pull_request:
+  workflow_dispatch:
+
+concurrency:
+  group: ci-${{ github.workflow }}-${{ github.ref }}
+  cancel-in-progress: true
+
+jobs:
+  build:
+    runs-on: ubuntu-latest
+    steps:
+      - uses: actions/checkout@v4
+      - uses: astral-sh/setup-uv@v5
+        with:
+          enable-cache: true
+      - name: Sync
+        run: uv sync --frozen
+      # --strict turns warnings into failures, so a broken link or a page
+      # missing from the nav fails here rather than shipping.
+      - name: Build
+        run: uv run mkdocs build --strict

From b38dcd403343942ad54b831a87f95e9c51cb1d5e Mon Sep 17 00:00:00 2001
From: Gotham-Zolio <18781106300@163.com>
Date: Thu, 10 Sep 2026 16:08:36 -0400
Subject: [PATCH 4/6] docs: add a Protocol section

The boundary between the training server and the env client is the reason
this project exists, and the site never described it. A reader could learn
how to add an environment or a policy, but not what the two processes say
to each other, or that an env client need not be Python.

The page orients rather than duplicates: SPEC.md stays in plugrl-protocol,
next to the code it describes, so the two cannot drift. What is here is the
exchange diagram, the three rules a first implementation usually gets wrong
- strict alternation, feedback env sets that need not match the infer's,
and chunk-summed reward - the conformance server, and the relationship to
openpi.

Both languages. mkdocs build --strict passes; nav translations go from 10
elements to 11.

Co-Authored-By: Claude Opus 5 (1M context) 
---
 docs/protocol/index.md    | 95 +++++++++++++++++++++++++++++++++++++++
 docs/protocol/index.zh.md | 79 ++++++++++++++++++++++++++++++++
 mkdocs.yml                |  4 ++
 3 files changed, 178 insertions(+)
 create mode 100644 docs/protocol/index.md
 create mode 100644 docs/protocol/index.zh.md

diff --git a/docs/protocol/index.md b/docs/protocol/index.md
new file mode 100644
index 0000000..84efc19
--- /dev/null
+++ b/docs/protocol/index.md
@@ -0,0 +1,95 @@
+# The wire protocol
+
+PlugRL splits a training run across two processes. A **training server**
+holds the policy and the learning algorithm. An **env client** runs
+environments, asks for actions, and reports what happened. They talk over
+WebSocket, with msgpack on the wire.
+
+That boundary is the reason an environment stack and a training stack never
+have to share a Python environment — and the reason an env client does not
+have to be Python at all. A ROS node can be one. So can a robot's onboard
+C++ controller.
+
+!!! info "The specification lives in `plugrl-protocol`"
+
+    **[SPEC.md](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)**
+    is the normative document. It is kept next to the code it describes so
+    the two cannot drift apart, and this page only orients you.
+
+    It also names its own known defects, in boxes marked **Gap**. Those are
+    the honest part; read them before building on anything.
+
+## The exchange
+
+```
+client                                     server
+  |------------ WebSocket handshake ---------->|
+  |<---------------- metadata -----------------|   the server speaks first
+  |------------------ infer ------------------>|   observations
+  |<----------------- action ------------------|   an action chunk
+  |----------------- feedback ---------------->|   reward, done, next obs
+```
+
+Four message types — `metadata`, `infer`, `action`, `feedback` — carried as
+msgpack maps with a `message_type` field. Arrays travel as
+
+```
+{b"__ndarray__": true, b"data": , b"dtype": "|
+  |<---------------- metadata -----------------|   服务端先说话
+  |------------------ infer ------------------>|   观测
+  |<----------------- action ------------------|   一段动作块
+  |----------------- feedback ---------------->|   奖励、终止标志、下一观测
+```
+
+四种消息类型 —— `metadata`、`infer`、`action`、`feedback` —— 都是带
+`message_type` 字段的 msgpack map。数组的形式是
+
+```
+{b"__ndarray__": true, b"data": , b"dtype": "
Date: Thu, 10 Sep 2026 22:17:17 -0400
Subject: [PATCH 5/6] docs: put a quickstart on the front page that actually
 learns

The quickstart was the dummy policy, whose learn is a sleep. A reader
following the front page of this site could confirm two processes talk to
each other and nothing more - which is a connectivity check, not a
quickstart, and the page did not say so.

It is now FPO on HalfCheetah-v5, CPU only, which trains: episode return
climbs out of the -300s in a few minutes. The connectivity check is kept
below it, labelled as what it is.

It also states the setting that would otherwise cost someone an afternoon.
FPO learns when its rollout buffer fills or when the run ends, so at the
default buffer_size=983040 a run shorter than a million steps learns exactly
once, at the very end, and produces a single point rather than a curve.

And it stops recommending the Ray launcher flatly. That line said only "Use
plugrl-run-server-ray for Ray-based distributed launch"; the launcher needs
the dppo extra, builds its worker list from the local GPU count so a
multi-node cluster still sees one node, and its server speaks an older
dialect of the protocol. Those are now stated, with a pointer to the spec.

Both languages. mkdocs build --strict passes.

Co-Authored-By: Claude Opus 5 (1M context) 
---
 docs/index.md    | 57 +++++++++++++++++++++++++++++++++++++++++++-----
 docs/index.zh.md | 47 ++++++++++++++++++++++++++++++++++++---
 2 files changed, 96 insertions(+), 8 deletions(-)

diff --git a/docs/index.md b/docs/index.md
index 49925c0..3841d05 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -6,17 +6,50 @@ PlugRL is an RL infrastructure for distributed experiments with a clean split be
 
 ## Quickstart
 
-Run a connectivity smoke test.
+Two processes: a training server that holds the policy, and an env client
+that runs environments and asks it for actions. This pair actually learns -
+FPO on HalfCheetah-v5, CPU only, no GPU and no assets to download.
+
+```bash
+# Terminal 1 - the training server
+plugrl-run-server fpo-policy default fpo default \
+    --port 8000 --policy.device cpu \
+    --algo.global-steps 500000 --algo.buffer-size 4096
+
+# Terminal 2 - the environment
+plugrl-run-env-client mujoco-v1 \
+    --server-host 127.0.0.1 --server-port 8000 \
+    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
+```
+
+Episode return climbs out of the -300s within a few minutes. `HalfCheetah-v5`
+has a 17-dimensional observation and a 6-dimensional action, which are
+exactly `fpo-policy`'s defaults, so nothing needs configuring. The
+environment needs `plugrl-env-client[mujoco]`.
+
+!!! warning "`--algo.buffer-size` is not decoration"
+
+    FPO learns when its rollout buffer fills, or when the run reaches its
+    last step. At the default `buffer_size=983040`, a run shorter than about
+    a million steps therefore learns **exactly once, at the very end** -
+    which gives you a single point instead of a curve.
+
+### Just checking connectivity?
 
 ```bash
 plugrl-run-server dummy-policy default dummy default
 plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
 ```
 
+The dummy algorithm's `learn` is a sleep - it moves no weights. Use it to
+confirm the two sides talk to each other, not to train anything.
+
 ## Verify
 
-- Server prints a WebSocket listening address.
-- Env client prints server metadata and starts stepping episodes.
+- The server prints a WebSocket listening address.
+- The env client prints the server's metadata - policy name, action shape -
+  and starts stepping episodes.
+- With `fpo`, the server prints a metrics table whose `rollout/reward` rises.
 
 ## Components
 
@@ -24,15 +57,29 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server
 - `plugrl-env-client`: environment runner, collects rollouts
 - `plugrl-protocol`: transport, message types, and serialization (WebSocket + msgpack)
 
+The boundary between the first two is [the protocol](protocol/index.md), and
+it is specified rather than implied: an env client does not have to be
+Python, or be this codebase.
+
 ## Common options
 
-- Env client connects to server via `--server-host` and `--server-port`.
-- Use `plugrl-run-server-ray` for Ray-based distributed launch.
+- The env client connects to the server via `--server-host` and `--server-port`.
+- `--num-procs` runs several env client processes against one server.
+
+!!! note "On `plugrl-run-server-ray`"
+
+    There is a Ray-based launcher, but it is **not a supported path today**.
+    It requires the `dppo` extra, builds its worker list from the *local*
+    GPU count so a multi-node cluster still only sees the head node, and its
+    server speaks an older dialect of the protocol than the WebSocket one -
+    see [SPEC.md section 5.3](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md).
+    Use `plugrl-run-server` unless you are working on the Ray path itself.
 
 ## Next steps
 
 - [User Guide](user_guide/index.md)
 - [Get Started](user_guide/get_started.md)
+- [Protocol](protocol/index.md)
 - [Algorithms](algorithm/index.md)
 - [Environments](env/index.md)
 - [Policies](policy/index.md)
diff --git a/docs/index.zh.md b/docs/index.zh.md
index 1d4edd2..4661db3 100644
--- a/docs/index.zh.md
+++ b/docs/index.zh.md
@@ -6,17 +6,46 @@ PlugRL 是一套面向分布式强化学习实验的基础设施。训练端与
 
 ## 快速开始
 
-先跑通一次联通性 smoke test。
+两个进程:训练端持有策略,环境端跑环境并向它请求动作。下面这一对**真的会学** ——
+FPO + HalfCheetah-v5,纯 CPU,不需要 GPU,也不需要下载任何资源文件。
+
+```bash
+# 终端 1 —— 训练端
+plugrl-run-server fpo-policy default fpo default \
+    --port 8000 --policy.device cpu \
+    --algo.global-steps 500000 --algo.buffer-size 4096
+
+# 终端 2 —— 环境端
+plugrl-run-env-client mujoco-v1 \
+    --server-host 127.0.0.1 --server-port 8000 \
+    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
+```
+
+几分钟内 episode 回报就会从 -300 附近爬上来。`HalfCheetah-v5` 的观测是 17 维、
+动作是 6 维,**正好是 `fpo-policy` 的默认值**,所以不需要任何配置。
+环境端需要 `plugrl-env-client[mujoco]`。
+
+!!! warning "`--algo.buffer-size` 不是装饰"
+
+    FPO 在 rollout buffer 填满时、或运行到最后一步时才学习。按默认的
+    `buffer_size=983040`,任何少于约一百万步的运行**只会在最后学一次** ——
+    你得到的是一个点,不是一条曲线。
+
+### 只想确认能连通?
 
 ```bash
 plugrl-run-server dummy-policy default dummy default
 plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
 ```
 
+dummy 算法的 `learn` 是一个 sleep,不会移动任何权重。它用来确认两端能对话,
+不是用来训练的。
+
 ## 验证
 
 - server 打印 WebSocket 监听地址
-- env client 打印 server 元信息并开始跑 episode
+- env client 打印 server 元信息(策略名、动作形状)并开始跑 episode
+- 用 `fpo` 时,server 的指标表里 `rollout/reward` 会上升
 
 ## 组件
 
@@ -24,15 +53,27 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server
 - `plugrl-env-client`:环境端,负责创建环境并采集 rollout
 - `plugrl-protocol`:协议与序列化层,WebSocket 与 msgpack
 
+前两者之间的边界就是[通信协议](protocol/index.zh.md),而且它是**被写下来的**而非
+默认的:环境端不必是 Python,也不必是这个代码库。
+
 ## 常用参数
 
 - env client 通过 `--server-host` 与 `--server-port` 连接 server
-- 分布式启动使用 `plugrl-run-server-ray`
+- `--num-procs` 可以起多个环境端进程连同一个 server
+
+!!! note "关于 `plugrl-run-server-ray`"
+
+    确实有一个基于 Ray 的启动器,但**目前不是受支持的路径**。它需要 `dppo`
+    extra;它用**本机**的 GPU 数构建 worker 列表,所以即使 Ray 连上多节点集群
+    也只看得见头节点;而且它的服务端说的是比 WebSocket 版更旧的协议方言,
+    见 [SPEC.md 第 5.3 节](https://github.com/PlugRL/plugrl-protocol/blob/main/SPEC.md)。
+    除非你就是在改 Ray 这条路径,否则请用 `plugrl-run-server`。
 
 ## 下一步
 
 - [用户指南](user_guide/index.zh.md)
 - [快速开始](user_guide/get_started.zh.md)
+- [通信协议](protocol/index.zh.md)
 - [算法](algorithm/index.zh.md)
 - [环境](env/index.zh.md)
 - [策略](policy/index.zh.md)

From 5d5cda8e5ebf991435e8165206debb874df8de8e Mon Sep 17 00:00:00 2001
From: Gotham-Zolio <18781106300@163.com>
Date: Thu, 10 Sep 2026 22:45:51 -0400
Subject: [PATCH 6/6] docs: a Get Started page that starts with installing

It had no installation section at all. It opened with `plugrl-run-server`,
which requires two packages that are not on PyPI and that the page never
mentioned cloning. A reader arriving from the front page had nowhere to go.

It now begins with the clone-and-uv-sync for both repositories, then the FPO
quickstart that actually learns, then the dummy connectivity check labelled
as what it is.

It also names the two flags that are not optional and previously were not
mentioned anywhere: --policy.device cpu, because the default is cuda and the
server dies on startup without a GPU, and --algo.buffer-size, because at the
default a short run learns once at the very end.

The troubleshooting section now covers what a new user actually hits,
including both of those.

The Ray launcher is qualified here and in the user guide rather than listed
as an equal alternative to plugrl-run-server.

Both languages. mkdocs build --strict passes.

Co-Authored-By: Claude Opus 5 (1M context) 
---
 docs/user_guide/get_started.md    | 97 +++++++++++++++++++++++++------
 docs/user_guide/get_started.zh.md | 95 ++++++++++++++++++++++++------
 docs/user_guide/index.md          | 11 +++-
 docs/user_guide/index.zh.md       | 10 +++-
 4 files changed, 170 insertions(+), 43 deletions(-)

diff --git a/docs/user_guide/get_started.md b/docs/user_guide/get_started.md
index 4387918..4198405 100644
--- a/docs/user_guide/get_started.md
+++ b/docs/user_guide/get_started.md
@@ -1,56 +1,117 @@
 # Get Started
 
-Run a minimal end-to-end smoke test.
+From nothing to a policy that is learning, in two terminals.
 
-> Note: Server default address is `0.0.0.0:8000`. Env client connects via `--server-host` and `--server-port`.
+## Install
+
+Neither package is on PyPI. Clone both and install each with `uv`:
+
+```bash
+git clone git@github.com:PlugRL/plugrl-server.git
+git clone git@github.com:PlugRL/plugrl-env-client.git
+
+cd plugrl-server     && uv sync && cd ..
+cd plugrl-env-client && uv sync --extra mujoco && cd ..
+```
+
+`--extra mujoco` is what the quickstart below needs. The env client has one
+extra per environment family; install only the ones you use.
+
+They can also live in one environment - the two dependency sets do coexist,
+which is measured in `plugrl-server/experiments/e1-dependency-conflict/`.
+Separate environments are simply the point of the split.
 
 ## Quickstart
 
-Terminal A starts the server.
+Terminal A, the training server:
 
 ```bash
-plugrl-run-server dummy-policy default dummy default
+plugrl-run-server fpo-policy default fpo default \
+    --port 8000 --policy.device cpu \
+    --algo.global-steps 500000 --algo.buffer-size 4096
 ```
 
-Terminal B starts the env client.
+Terminal B, the environment:
 
 ```bash
-plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
+plugrl-run-env-client mujoco-v1 \
+    --server-host 127.0.0.1 --server-port 8000 \
+    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
 ```
 
+`HalfCheetah-v5` is 17 observation dimensions and 6 action dimensions, which
+are `fpo-policy`'s defaults, so nothing needs configuring.
+
+!!! warning "Two settings that are not decoration"
+
+    `--policy.device cpu` - the default is `cuda`, and without a GPU the
+    server fails on startup with `Torch not compiled with CUDA enabled`.
+
+    `--algo.buffer-size 4096` - FPO learns when its rollout buffer fills or
+    when the run reaches its last step. At the default `983040`, a run
+    shorter than about a million steps learns **once, at the very end**,
+    giving a single point instead of a curve.
+
 ## Verify
 
-- Server prints a WebSocket listening address.
-- Env client prints server metadata and steps episodes.
+- The server prints a WebSocket listening address.
+- The env client prints the server's metadata, including `action_dim` and
+  `action_horizon`, and starts stepping episodes.
+- The server's metrics show `rollout/reward` rising. On HalfCheetah it
+  starts near -300 and climbs out within a few minutes.
 
-## DPPO examples
+`plugrl-server/experiments/e6-first-learning-curve/` holds a three-seed run
+of exactly this, with the script that produced it.
 
-Single process server.
+## Just checking connectivity
 
 ```bash
-plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp
+plugrl-run-server dummy-policy default dummy default
+plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
 ```
 
-Ray launcher.
+The dummy algorithm's `learn` is a sleep and moves no weights. Use it to
+confirm the two sides talk, not to train.
+
+## Other policies
 
 ```bash
-plugrl-run-server-ray dppo-policy default dppo hopper --exp_name my_dppo_exp --num-ddp-gpus 4
+plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp
 ```
 
+DPPO needs `plugrl-server[dppo]` and a pretrained checkpoint. `pi0-policy`
+needs a checkpoint too, and a GPU.
+
+!!! note "The Ray launcher is not a supported path today"
+
+    `plugrl-run-server-ray` exists, but it requires the `dppo` extra, builds
+    its worker list from the *local* GPU count - so a multi-node cluster
+    still only sees the head node - and its server speaks an older dialect of
+    the protocol than the WebSocket one. Use `plugrl-run-server` unless you
+    are working on the Ray path itself.
+
 ## Common options
 
-- Set server address with `--host` and `--port`.
-- Control episode count with `--num-episodes`.
-- Resume requires an existing experiment directory under `--checkpoint-base-dir`.
+- Set the server address with `--host` and `--port`.
+- Control the episode count with `--num-episodes`.
+- `--num-procs` runs several env client processes against one server.
+- `--resume` requires an existing experiment directory under
+  `--checkpoint-base-dir`.
 
 ## Troubleshooting
 
-- Env client keeps retrying: check server address and firewall.
-- `--resume` requires existing checkpoints.
+| Symptom | Cause |
+|---|---|
+| Env client keeps retrying | The server is not listening yet, or the address or firewall is wrong |
+| `Torch not compiled with CUDA enabled` | Pass `--policy.device cpu` |
+| Only one metrics row, at the very end | `--algo.buffer-size` is larger than the run |
+| `--resume` raises `FileNotFoundError` | There are no checkpoints in that directory yet |
+| An environment reports a missing extra | Install it, e.g. `uv sync --extra mujoco` |
 
 ## Next steps
 
 - [User Guide](index.md)
+- [Protocol](../protocol/index.md)
 - [Algorithms](../algorithm/index.md)
 - [Environments](../env/index.md)
 - [Policies](../policy/index.md)
diff --git a/docs/user_guide/get_started.zh.md b/docs/user_guide/get_started.zh.md
index fba7a79..bde10b9 100644
--- a/docs/user_guide/get_started.zh.md
+++ b/docs/user_guide/get_started.zh.md
@@ -1,56 +1,113 @@
 # 快速开始
 
-用最小组件跑通端到端 smoke test。
+从零到一个正在学习的策略,两个终端。
 
-> Note: server 默认地址为 `0.0.0.0:8000`。env client 通过 `--server-host` 与 `--server-port` 连接。
+## 安装
 
-## 快速开始命令
+两个包都不在 PyPI 上。分别 clone 并用 `uv` 安装:
 
-终端 A 启动 server。
+```bash
+git clone git@github.com:PlugRL/plugrl-server.git
+git clone git@github.com:PlugRL/plugrl-env-client.git
+
+cd plugrl-server     && uv sync && cd ..
+cd plugrl-env-client && uv sync --extra mujoco && cd ..
+```
+
+`--extra mujoco` 是下面快速开始所需的。env client 的每个环境家族对应一个 extra,
+只装你要用的即可。
+
+两者也可以装在同一个环境里 —— 这两套依赖确实能共存,
+`plugrl-server/experiments/e1-dependency-conflict/` 里有实测。
+分开装只是这个架构的本意。
+
+## 快速开始
+
+终端 A,训练端:
 
 ```bash
-plugrl-run-server dummy-policy default dummy default
+plugrl-run-server fpo-policy default fpo default \
+    --port 8000 --policy.device cpu \
+    --algo.global-steps 500000 --algo.buffer-size 4096
 ```
 
-终端 B 启动 env client。
+终端 B,环境端:
 
 ```bash
-plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
+plugrl-run-env-client mujoco-v1 \
+    --server-host 127.0.0.1 --server-port 8000 \
+    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
 ```
 
+`HalfCheetah-v5` 是 17 维观测、6 维动作,**正好是 `fpo-policy` 的默认值**,
+所以不需要任何配置。
+
+!!! warning "两个不是装饰的参数"
+
+    `--policy.device cpu` —— 默认是 `cuda`,没有 GPU 时服务端会在启动时
+    直接报 `Torch not compiled with CUDA enabled`。
+
+    `--algo.buffer-size 4096` —— FPO 在 rollout buffer 填满时、或运行到最后
+    一步时才学习。按默认的 `983040`,少于约一百万步的运行**只会在最后学一次**,
+    你得到的是一个点而不是一条曲线。
+
 ## 验证
 
 - server 打印 WebSocket 监听地址
-- env client 打印 server 元信息并开始跑 episode
+- env client 打印 server 元信息(含 `action_dim`、`action_horizon`)并开始跑 episode
+- server 的指标里 `rollout/reward` 在上升。HalfCheetah 上从 -300 附近起步,
+  几分钟内就会爬上来
 
-## DPPO 示例
+`plugrl-server/experiments/e6-first-learning-curve/` 里有一次三种子的完整运行,
+以及产生它的脚本。
 
-单进程 server。
+## 只想确认能连通
 
 ```bash
-plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp
+plugrl-run-server dummy-policy default dummy default
+plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
 ```
 
-Ray 启动。
+dummy 算法的 `learn` 是一个 sleep,不移动任何权重。它用来确认两端能对话,
+不是用来训练的。
+
+## 其他策略
 
 ```bash
-plugrl-run-server-ray dppo-policy default dppo hopper --exp_name my_dppo_exp --num-ddp-gpus 4
+plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp
 ```
 
+DPPO 需要 `plugrl-server[dppo]` 和一个预训练 checkpoint。`pi0-policy` 同样需要
+checkpoint,而且需要 GPU。
+
+!!! note "Ray 启动器目前不是受支持的路径"
+
+    `plugrl-run-server-ray` 确实存在,但它需要 `dppo` extra;它用**本机**的
+    GPU 数构建 worker 列表,所以即使连上多节点集群也只看得见头节点;
+    而且它的服务端说的是比 WebSocket 版更旧的协议方言。
+    除非你就是在改 Ray 这条路径,否则请用 `plugrl-run-server`。
+
 ## 常用参数
 
-- server 地址使用 `--host` 与 `--port` 修改
-- 运行 episode 数使用 `--num-episodes`
-- resume 需要 `--checkpoint-base-dir` 下存在实验目录与 checkpoint
+- 用 `--host` 和 `--port` 设置服务端地址
+- 用 `--num-episodes` 控制 episode 数
+- `--num-procs` 可以起多个环境端进程连同一个 server
+- `--resume` 需要 `--checkpoint-base-dir` 下已有实验目录
 
-## 常见问题
+## 排错
 
-- env client 一直重试:检查 server 是否已启动,host 与 port 是否一致,端口是否可达。
-- `--resume` 找不到 checkpoint:确认实验目录存在且包含 checkpoint。
+| 现象 | 原因 |
+|---|---|
+| env client 一直重试 | server 还没监听,或地址/防火墙不对 |
+| `Torch not compiled with CUDA enabled` | 加上 `--policy.device cpu` |
+| 指标只有最后一行 | `--algo.buffer-size` 比整个运行还大 |
+| `--resume` 报 `FileNotFoundError` | 那个目录下还没有任何 checkpoint |
+| 某个环境报缺少 extra | 装上它,例如 `uv sync --extra mujoco` |
 
 ## 下一步
 
 - [用户指南](index.zh.md)
+- [通信协议](../protocol/index.zh.md)
 - [算法](../algorithm/index.zh.md)
 - [环境](../env/index.zh.md)
 - [策略](../policy/index.zh.md)
diff --git a/docs/user_guide/index.md b/docs/user_guide/index.md
index 5f5d7a0..014b78e 100644
--- a/docs/user_guide/index.md
+++ b/docs/user_guide/index.md
@@ -5,10 +5,13 @@ Run PlugRL end to end: start a server, then start one or more env clients.
 ## Quickstart
 
 ```bash
-plugrl-run-server dummy-policy default dummy default
-plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
+plugrl-run-server fpo-policy default fpo default \n    --policy.device cpu --algo.global-steps 500000 --algo.buffer-size 4096
+plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \n    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
 ```
 
+That pair learns. [Get Started](get_started.md) explains the two flags that
+are not optional, and has the dummy connectivity check.
+
 ## Verify
 
 - Server prints a WebSocket listening address.
@@ -16,7 +19,9 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server
 
 ## Workflow
 
-1. Start a training server with `plugrl-run-server` or `plugrl-run-server-ray`.
+1. Start a training server with `plugrl-run-server`. (There is also
+   `plugrl-run-server-ray`, but it is not a supported path today - see
+   [Get Started](get_started.md).)
 2. Start one or more env clients with `plugrl-run-env-client `.
 
 ## Components
diff --git a/docs/user_guide/index.zh.md b/docs/user_guide/index.zh.md
index 1ebe4c4..d81b2b0 100644
--- a/docs/user_guide/index.zh.md
+++ b/docs/user_guide/index.zh.md
@@ -5,10 +5,13 @@
 ## 快速开始
 
 ```bash
-plugrl-run-server dummy-policy default dummy default
-plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server-port 8000
+plugrl-run-server fpo-policy default fpo default \n    --policy.device cpu --algo.global-steps 500000 --algo.buffer-size 4096
+plugrl-run-env-client mujoco-v1 --server-host 127.0.0.1 --server-port 8000 \n    --num-envs 1 --num-episodes 600 --runner.replan-steps 1 --runner.seed 0
 ```
 
+这一对**真的会学**。[快速开始](get_started.zh.md)里说明了那两个不可省的参数,
+以及 dummy 连通性检查怎么做。
+
 ## 验证
 
 - server 打印 WebSocket 监听地址
@@ -16,7 +19,8 @@ plugrl-run-env-client dummy-v1 --num-episodes 2 --server-host 127.0.0.1 --server
 
 ## 流程
 
-1. 用 `plugrl-run-server` 或 `plugrl-run-server-ray` 启动训练端。
+1. 用 `plugrl-run-server` 启动训练端。(也有 `plugrl-run-server-ray`,但它目前
+   不是受支持的路径,见[快速开始](get_started.zh.md))
 2. 用 `plugrl-run-env-client ` 启动一个或多个环境端。
 
 ## 组件