diff --git a/docs/index.md b/docs/index.md index 53ddfe4..a8a14be 100644 --- a/docs/index.md +++ b/docs/index.md @@ -44,6 +44,15 @@ MuJoCo tasks, and robosuite 1.4.1 with MuJoCo 2.3.7 for robomimic, because robosuite 1.4.1 does not run on MuJoCo 3. One server codebase trained all sixteen. +The same pair also learns with its env clients on another machine: a +Windows laptop on campus Wi-Fi steps HalfCheetah for a server on a Linux +workstation. Six seeds on each side fall within one band. Across the two +machines, a run took 41-53 minutes instead of 20 +([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training)). + +Episode return over 100 iterations for fpo-policy with FPO on HalfCheetah, six seeds on one machine and six across two machines. Both bands rise from about -300 to between roughly 700 and 2,300 and overlap throughout. + ## The environment side is light The training server is 6.5G and wants a GPU. The machine running environments @@ -57,18 +66,26 @@ Stepping is 10x slower in software, and most of that hides behind the queue of clients waiting on one policy. It does not have to be Python either. [The protocol](protocol/index.md) is -written down, with a conformance checker, and a client in C++ with nothing -beyond the standard library - no msgpack or WebSocket library - drove a real -training server through 120 training exchanges -([E2](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e2-cross-language)). +written down, with a conformance checker. A C++ program with nothing beyond +the standard library - no msgpack or WebSocket library - steps its own copy of +Pendulum and trains a policy through it, as well as the Python env client does +([E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum)). ## What the split costs -Little. With a 184 KiB observation, one exchange takes about 0.8 ms on one -machine, and leaving the machine adds about 0.5 ms -([E7](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e7-cross-machine)). -That second number was measured from a virtual machine to its host; it has -not yet been measured between two physical machines. +On one machine, almost nothing: an exchange takes 0.1 ms with states only and +under 1 ms with a 588 KiB observation. Between two machines it becomes two +terms: +- **A fixed latency.** About 3 ms between a laptop on campus Wi-Fi and a wired + workstation, over Tailscale. +- **Twice the observation's bytes over the link's bandwidth.** Each exchange + carries the observation twice, once asking for the action and once in the + feedback. + +On that link, at 18 MB/s, a state-only task pays about 3 ms per step, and a +184 KiB camera observation about 21 ms +([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training)). +A faster link shrinks the second term in proportion. ## A real VLA, end to end diff --git a/docs/index.zh.md b/docs/index.zh.md index aa01281..d19dcdf 100644 --- a/docs/index.zh.md +++ b/docs/index.zh.md @@ -32,6 +32,13 @@ robosuite 或 gymnasium。环境客户端分别跑在两套独立的环境里: gymnasium 加 MuJoCo 3;robomimic 用的是 robosuite 1.4.1 加 MuJoCo 2.3.7,因为 robosuite 1.4.1 在 MuJoCo 3 上跑不起来。训练这十六格的是同一个服务端代码库。 +同样这一对,把环境端放到另一台机器上也照样学会:一台连着校园 Wi-Fi 的 Windows 笔记本步进 +HalfCheetah,训练端在一台 Linux 工作站上。两边各六个种子落在同一条带子里;跨两台机器时一次 +训练用了 41 到 53 分钟,同一台机器上是 20 分钟([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training))。 + +fpo-policy 用 FPO 训练 HalfCheetah 100 轮的回报:一台机器上六个种子、跨两台机器六个种子,两条带子都从约 -300 升到约 700 至 2300 之间,全程重叠。 + ## 环境端很轻 训练端有 6.5G,而且要一张显卡;跑环境的那台机器两样都不需要。LIBERO 环境端可以装成 @@ -41,16 +48,20 @@ robosuite 1.4.1 在 MuJoCo 3 上跑不起来。训练这十六格的是同一个 [E13](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e13-gpu-free-rendering))。 软件渲染的单步慢 10 倍,其中大部分被"多个客户端排队等同一个策略"的等待掩盖了。 -环境端也不必是 Python。[协议](protocol/index.zh.md)是写下来的,附带一致性检查器;一个 -除了标准库之外什么都不用的 C++ 客户端(没有 msgpack 库,也没有 WebSocket 库),驱动一个 -真实的训练端完成了 120 次训练交换 -([E2](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e2-cross-language))。 +环境端也不必是 Python。[协议](protocol/index.zh.md)是写下来的,附带一致性检查器。一个 +除了标准库之外什么都不用的 C++ 程序(没有 msgpack 库,也没有 WebSocket 库),自己实现了 +Pendulum,通过 PlugRL 训练出了策略,学得和 Python 环境端一样好 +([E44](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e44-cpp-pendulum))。 ## 拆开的代价 -很小。观测为 184 KiB 时,同一台机器上一次交换约 0.8 毫秒,离开这台机器再多约 0.5 毫秒 -([E7](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e7-cross-machine))。 -后一个数是从虚拟机到它的宿主机测的,还没有在两台物理机之间测过。 +同一台机器上几乎没有代价:只传状态时一次交换 0.1 毫秒,观测有 588 KiB 也不到 1 毫秒。 +跨两台机器时,代价分成两项: +- **固定延迟**:校园 Wi-Fi 上的笔记本经 Tailscale 连有线工作站,大约 3 毫秒。 +- **两倍的观测数据量除以链路带宽**:每次交换里观测要传两次,一次在请求动作时,一次在反馈里。 + +在这条链路上(18 MB/s),只传状态的任务每步多约 3 毫秒,一个 184 KiB 的相机观测每步多约 +21 毫秒([E43](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e43-cross-machine-training))。链路越快,第二项按比例越小。 ## 真实 VLA,端到端 diff --git a/docs/media/cross-machine.png b/docs/media/cross-machine.png new file mode 100644 index 0000000..980508f Binary files /dev/null and b/docs/media/cross-machine.png differ