Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 26 additions & 9 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,15 @@ MuJoCo tasks, and robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic, because
robosuite 1.4.1 does not run on MuJoCo 3. One server codebase trained all
sixteen.

The same pair also learns with its env clients on another machine: a
Windows laptop on campus Wi-Fi steps HalfCheetah for a server on a Linux
workstation. Six seeds on each side fall within one band. Across the two
machines, a run took 41-53 minutes instead of 20
([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training)).

<img src="/media/cross-machine.png" style="width:100%;max-width:720px"
alt="Episode return over 100 iterations for fpo-policy with FPO on HalfCheetah, six seeds on one machine and six across two machines. Both bands rise from about -300 to between roughly 700 and 2,300 and overlap throughout.">

## The environment side is light

The training server is 6.5G and wants a GPU. The machine running environments
Expand All @@ -57,18 +66,26 @@ Stepping is 10x slower in software, and most of that hides behind the queue
of clients waiting on one policy.

It does not have to be Python either. [The protocol](protocol/index.md) is
written down, with a conformance checker, and a client in C++ with nothing
beyond the standard library - no msgpack or WebSocket library - drove a real
training server through 120 training exchanges
([E2](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e2-cross-language)).
written down, with a conformance checker. A C++ program with nothing beyond
the standard library - no msgpack or WebSocket library - steps its own copy of
Pendulum and trains a policy through it, as well as the Python env client does
([E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum)).

## What the split costs

Little. With a 184 KiB observation, one exchange takes about 0.8 ms on one
machine, and leaving the machine adds about 0.5 ms
([E7](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e7-cross-machine)).
That second number was measured from a virtual machine to its host; it has
not yet been measured between two physical machines.
On one machine, almost nothing: an exchange takes 0.1 ms with states only and
under 1 ms with a 588 KiB observation. Between two machines it becomes two
terms:
- **A fixed latency.** About 3 ms between a laptop on campus Wi-Fi and a wired
workstation, over Tailscale.
- **Twice the observation's bytes over the link's bandwidth.** Each exchange
carries the observation twice, once asking for the action and once in the
feedback.

On that link, at 18 MB/s, a state-only task pays about 3 ms per step, and a
184 KiB camera observation about 21 ms
([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training)).
A faster link shrinks the second term in proportion.

## A real VLA, end to end

Expand Down
25 changes: 18 additions & 7 deletions docs/index.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,13 @@ robosuite 或 gymnasium。环境客户端分别跑在两套独立的环境里:
gymnasium 加 MuJoCo 3;robomimic 用的是 robosuite 1.4.1 加 MuJoCo 2.3.7,因为
robosuite 1.4.1 在 MuJoCo 3 上跑不起来。训练这十六格的是同一个服务端代码库。

同样这一对,把环境端放到另一台机器上也照样学会:一台连着校园 Wi-Fi 的 Windows 笔记本步进
HalfCheetah,训练端在一台 Linux 工作站上。两边各六个种子落在同一条带子里;跨两台机器时一次
训练用了 41 到 53 分钟,同一台机器上是 20 分钟([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training))。

<img src="/media/cross-machine.png" style="width:100%;max-width:720px"
alt="fpo-policy 用 FPO 训练 HalfCheetah 100 轮的回报:一台机器上六个种子、跨两台机器六个种子,两条带子都从约 -300 升到约 700 至 2300 之间,全程重叠。">

## 环境端很轻

训练端有 6.5G,而且要一张显卡;跑环境的那台机器两样都不需要。LIBERO 环境端可以装成
Expand All @@ -41,16 +48,20 @@ robosuite 1.4.1 在 MuJoCo 3 上跑不起来。训练这十六格的是同一个
[E13](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e13-gpu-free-rendering))。
软件渲染的单步慢 10 倍,其中大部分被"多个客户端排队等同一个策略"的等待掩盖了。

环境端也不必是 Python。[协议](protocol/index.zh.md)是写下来的,附带一致性检查器;一个
除了标准库之外什么都不用的 C++ 客户端(没有 msgpack 库,也没有 WebSocket 库),驱动一个
真实的训练端完成了 120 次训练交换
([E2](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e2-cross-language))。
环境端也不必是 Python。[协议](protocol/index.zh.md)是写下来的,附带一致性检查器。一个
除了标准库之外什么都不用的 C++ 程序(没有 msgpack 库,也没有 WebSocket 库),自己实现了
Pendulum,通过 PlugRL 训练出了策略,学得和 Python 环境端一样好
([E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum))。

## 拆开的代价

很小。观测为 184 KiB 时,同一台机器上一次交换约 0.8 毫秒,离开这台机器再多约 0.5 毫秒
([E7](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e7-cross-machine))。
后一个数是从虚拟机到它的宿主机测的,还没有在两台物理机之间测过。
同一台机器上几乎没有代价:只传状态时一次交换 0.1 毫秒,观测有 588 KiB 也不到 1 毫秒。
跨两台机器时,代价分成两项:
- **固定延迟**:校园 Wi-Fi 上的笔记本经 Tailscale 连有线工作站,大约 3 毫秒。
- **两倍的观测数据量除以链路带宽**:每次交换里观测要传两次,一次在请求动作时,一次在反馈里。

在这条链路上(18 MB/s),只传状态的任务每步多约 3 毫秒,一个 184 KiB 的相机观测每步多约
21 毫秒([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training))。链路越快,第二项按比例越小。

## 真实 VLA,端到端

Expand Down
Binary file added docs/media/cross-machine.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading