Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,14 +24,15 @@ has been measured:
| [coverage](#what-runs-on-it) | Which policies and algorithms learn through the split? | **Every one tried** - sixteen combinations on four tasks, all learning, below |
| [E2](experiments/e2-cross-language/) | Does an env client have to be this codebase, or Python? | **No** - a C++ client with no third-party libraries drove this server |
| [E12](experiments/e12-cuda-free-rollout/) | Does a rollout machine need CUDA, or a GPU? | **No** - LIBERO's env client goes from 7.8G to 3.4G with no nvidia wheels, and renders on the CPU at 1.91x the wall clock ([E13](experiments/e13-gpu-free-rendering/)) |
| [E7](experiments/e7-cross-machine/) | What does the boundary cost once packets leave the machine? | **+0.52 ms** on a 184 KiB observation, measured from a VM to its host |
| [E10](experiments/e10-vla-forward-cost/) | Is that cheap beside a VLA forward pass? | **Yes** - 1.3-3.6% of a step |
| [E43](experiments/e43-cross-machine-training/) | Does training still work with the env clients on another physical machine? | **Yes** - the quickstart pair learns with its env clients on a Windows laptop over campus Wi-Fi |
| [E43](experiments/e43-cross-machine-training/) | What does crossing cost? | About **3 ms plus twice the observation's bytes over the link** per exchange: 21 ms for 184 KiB at 18 MB/s |
| [E10](experiments/e10-vla-forward-cost/) | Is that cheap beside a VLA forward pass? | **On a fast link.** Over that Wi-Fi a 184 KiB observation is 21% of pi0.5's 100 ms forward, not E10's 1.3-3.6% |
| [E11](experiments/e11-vla-rl-libero/) | Does a full-size VLA go through it? | **Yes** - pi0.5 on LIBERO scores what openpi publishes, and the server's episode and step counts match the clients' exactly |

Fine-tuning pi0.5 with reinforcement learning through it has not made the
policy better yet; that record is on
[its own page](https://plugrl.github.io/vla/).
[`experiments/`](experiments/) holds forty experiments, thirty-nine of them
[`experiments/`](experiments/) holds forty-one experiments, forty of them
run. Each carries its data and a `FINDINGS.md` stating what the result does
**not** support. The documentation, including a quickstart that trains FPO
on HalfCheetah with no GPU and nothing to download, is at
Expand Down
10 changes: 6 additions & 4 deletions experiments/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,10 +173,12 @@ rather than an achievement.

## What is missing

- **A true cross-machine number.** E7 got as far as one virtual machine to
its host, which separates the network stack from the machine boundary but
still shares a CPU and a hypervisor. A physical NIC and a switch will cost
more; how much more is unmeasured, and needs a second computer.
- **A cross-machine number on a fast link.** E7 got as far as one virtual
machine to its host. E43 then crossed two physical machines - a wired Linux
workstation and a Windows laptop on campus Wi-Fi - and found the cost is a
fixed latency (about 3 ms) plus twice the observation's bytes over the
link's bandwidth (18 MB/s there), with training unchanged. A wired LAN or a
datacenter link is still unmeasured; the model says how it scales.
- **E3 was never run.** Its protocol is pre-registered, including an explicit
declaration of the familiarity bias that would have favoured PlugRL and
three ranked mitigations, but no data exists.
Expand Down
32 changes: 32 additions & 0 deletions experiments/e43-cross-machine-training/AMENDMENT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# E43 amendments

## 1 - 2026-09-28, after the registered run: three more seeds per arm

**What prompted it.** The registered run holds on every prediction. But the
two arms ended apart:
- The cross arm's last ten iterations averaged 2,060, 2,127 and 946.
- The local arm's averaged 1,262, 695 and 725.

The figure shows that gap from about iteration 40 on. PROTOCOL.md predicted
nothing about it. The two arms should draw from the same distribution:
- the same code
- the same settings
- the same seeds
- trajectories that differ only by floating-point bits (known item 4)

A real gap would then be a defect to find, not a result. With three seeds per
arm, luck cannot be told from a difference.

**What is added.** Seeds 3, 4 and 5 in both arms, with the same scripts, the
same concurrent layout and the same machines. They are written to
`results/local-b`, `results/cross-b` and `results/cross-clients-b`.

**How it is read.** The additions are reported, not predicted. P1-P6 stay as
registered, on seeds 0-2. The six seeds per arm are compared on:
- iteration 91-100's mean
- the status rule
- the extra time, as P4 measured it

**What would count as a gap worth chasing:** all six cross seeds above all
six local seeds, or the arms' six-seed means further apart than the spread
within either arm. Anything less is read as seed variation.
158 changes: 158 additions & 0 deletions experiments/e43-cross-machine-training/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,158 @@
# E43: with its env clients on another physical machine, the cell still learns; crossing costs a fixed latency plus the observation's bytes over the link

2026-09-28 · server on guangzhao (Linux workstation, wired); env clients on
a Windows 11 laptop on Wi-Fi 6, joined by Tailscale peer-to-peer on one
campus network · protocol: [`PROTOCOL.md`](PROTOCOL.md) (`46bce03`, after the
pilot in `results/pilot.txt`, before the registered run) · amendment:
[`AMENDMENT.md`](AMENDMENT.md)

---

## The result

**A. Training.** The coverage figure's `fpo-policy` · FPO · HalfCheetah
cell, with E16's command lines, ran as two arms at once. In the local arm,
servers and env clients were on guangzhao. In the cross arm, the servers
were on guangzhao and the env clients on the laptop. The table gives the mean
return over iterations 91-100 against iteration 1:

| seed | local | cross |
| --- | --- | --- |
| 0 | -293 -> 1,262 (**+1,555**) | -345 -> 2,060 (**+2,405**) |
| 1 | -325 -> 695 (**+1,020**) | -318 -> 2,127 (**+2,445**) |
| 2 | -251 -> 725 (**+976**) | -250 -> 946 (**+1,196**) |

* **P1 holds.** All six seeds logged 100 iterations and five checkpoints,
with no traceback, and every client exited 0.
* **V1 passes.** The cross clients connected to guangzhao's Tailscale
address and the local ones to 127.0.0.1. All six servers logged the same
configuration.
* **P2 and P3 hold.** Both arms learn, 3 of 3 seeds each, against a bar of
+200.
* **P4 holds.** The cross arm took 1,253, 1,288 and 1,356 s longer from
iteration 1 to 100, against a predicted 1,119 s ±25%. All three are inside
the band, but all three sit above the point prediction, by 12-21%.

**B. The boundary's cost**, as the median of five runs of 1,000 exchanges.
Throughput was 20.5, 18.0, 19.6, 15.9 and 14.4 MB/s (median 18.0).

| payload | `lo` | `ts` | `phys` | `phys` predicted |
| --- | --- | --- | --- | --- |
| states only | 0.099 ms | 0.107 ms | 2.82 ms | |
| 48 KiB | 0.161 | 0.172 | 8.01 | 8.28 |
| 184 KiB | 0.293 | 0.306 | 21.38 | 23.72 |
| 588 KiB | 0.947 | 0.700 | 65.80 | 69.72 |

* **V2 passes.** All 60 ladder runs completed.
* **P5 holds.** `lo` and `ts` agree within 0.25 ms at every payload. The
588 KiB row is the exception in kind, if not in verdict: 0.248 ms apart,
just inside the tolerance, and with `lo` the slower of the two - as in the
pilot. That is not explained.
* **P6 holds.** The crossing is a fixed latency plus twice the observation's
bytes over the link's throughput. Measured is 3-10% under predicted at
all three payloads.

---

## What crossing costs, and what it is made of

On one machine the boundary is a tenth of a millisecond with states only,
and under a millisecond at 588 KiB. Tailscale adds nothing measurable there:
a machine's own Tailscale address is short-circuited like any local one,
as E7 found for a machine's own Ethernet address.

Between two machines it becomes two terms:
- **A fixed one**: about 2.8 ms here, which is Wi-Fi plus WireGuard plus two
operating systems.
- **One proportional to size**: each exchange sends the observation twice,
once in the previous feedback and once in the infer, and both queue on the
same link. At 18 MB/s that is about 0.11 ms per KiB of observation, more
than fifty times E7's VM-to-host 2 µs.

A state-only task like this cell pays almost only the fixed term. Per call,
the env client waited 5.0 ms for an action across machines against 2.7 ms
locally. Stepping MuJoCo on the laptop took 0.50 ms against guangzhao's
0.10, and sending feedback 0.41 ms against 0.07. That doubled the wall clock,
from 20 minutes to 41, and changed nothing else; an hour later, with a busier
link and laptop, it was 52 minutes (below).

An image-based policy pays the second term. **This corrects E10**, whose
"1.3-3.6% of a step" used E7's 1.3 ms VM-to-host crossing. Over this link a
184 KiB observation costs 21.4 ms per exchange, which is 21% of the
full-size pi0.5's 100 ms forward. On a faster link the second term shrinks
in proportion to the bandwidth, but no faster link was measured here.

P4's three misses on the high side fit a link that was slowing during the
session: the five throughput runs fell from 20.5 to 14.4 MB/s. This is an
observation, not a test.

---

## The two arms' returns

The registered run left the cross arm's seeds higher: 2,060, 2,127 and 946,
against 1,262, 695 and 725. PROTOCOL.md did not predict that. The two arms
draw from one distribution, so a real gap would have been a defect to find.
[`AMENDMENT.md`](AMENDMENT.md) added seeds 3-5 to both arms, and fixed how
they would be read, before they ran. Iterations 91-100:

| seed | local | cross |
| --- | --- | --- |
| 0 | 1,262 | 2,060 |
| 1 | 695 | 2,127 |
| 2 | 725 | 946 |
| 3 | 1,967 | 2,267 |
| 4 | 2,336 | 1,587 |
| 5 | 857 | 1,082 |
| mean (sd) | 1,307 (694) | 1,678 (564) |

All twelve seeds learn by the status rule. Two of the new local seeds end
above four of the cross seeds, and the means are 371 apart, against a spread
of 694 within the local arm. By the amendment's rule this is seed variation
(`results/amendment.txt`). That is why the figure draws each arm's band, from
its lowest seed to its highest, and no mean line.

![Both arms' seeds within one band](results/cross-machine.png)

The amendment's cross arm also took longer: 1,875-1,985 s more than its local
seeds, against 1,253-1,356 s in the registered run. Two things were slower
an hour later.
- The link: the env clients waited 6.0 ms per call for an action instead of
5.0.
- The laptop, which was doing other work: stepping MuJoCo took 0.73 ms
instead of 0.50, and sending feedback 0.56 ms instead of 0.41.

Together that is about 1.4 ms per step, or 570 s over the run, which is the
difference. On a shared Wi-Fi link the crossing's cost is not a constant.

---

## What E43 does not show

* **Any other link, or this one at a steady cost.** Campus Wi-Fi through
Tailscale is slower than two machines on one wired switch, and it changed
within the evening (throughput 20.5 down to 14.4 MB/s; per-call wait 5.0
then 6.0 ms). The model says how the cost scales, but only this link was
measured.
* **Any other second machine**, or more than one. This is one laptop and
one server.
* **Image-based training across machines.** The cell sends states. Part B
measures what images cost on this link; no image policy was trained
across it.
* **Identical trajectories.** The two machines' MuJoCo differ in the last
bit from the first step (8.9e-16) and HalfCheetah amplifies it (order 1
by step 100, `results/pilot.txt`). On one machine a seed reproduces
exactly. So the arms are compared by the status rule, not trajectory by
trajectory.
* **Anything about an environment that does not wait.** PlugRL's env client
stops and waits for each action, so latency slows training without
changing its data. A robot, or any environment that keeps moving, turns
latency into a different problem, and E43 does not touch it.

The runs' checkpoints and tensorboards stay on guangzhao. Kept here:
- the logs, of both arms and of the amendment's seeds
- both machines' client summaries
- `ladder.tsv` and `throughput.tsv`
- `curves.json` and `cross-machine.png` (`figure.py`)
- `verdicts.txt` and `summary.tsv` (`summarise.py`)
- `amendment.txt` (`amendment.py`)
169 changes: 169 additions & 0 deletions experiments/e43-cross-machine-training/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,169 @@
# E43 measurement protocol (pre-registered)

**Written 2026-09-28, after the pilot in `results/pilot.txt`, before the
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The questions

PlugRL's home page says the env clients can run on another machine. Every
cell of the coverage figure was trained with both sides on one machine, and
E7's crossing was a VM to its own host. So:

**A. With the env clients on another physical machine, does training still
learn?**

**B. What does crossing to that machine cost per exchange, and what is the
cost made of?**

---

## Declared in advance: what was already known

1. **The cell.** `fpo-policy` · FPO · HalfCheetah-v5 from the coverage
figure, with E16 Phase A's command lines (`server_side.sh`): 409,600
steps, buffer 4,096, so 100 iterations. One server per seed, and the same
seed for the server and the client. Its status rule, as the figure has
it: it **learns** if the mean return over iterations 91-100 is at least
**+200** over iteration 1, on at least **2 of 3** seeds.
2. **The machines** (`results/pilot.txt`).
- The server runs on guangzhao, a wired Linux workstation.
- The second machine is a Windows 11 laptop on Wi-Fi 6. Its env client
runs natively on Windows, not in WSL.
- Both are on the same campus network and joined by Tailscale
peer-to-peer, not relayed. guangzhao's public address refuses inbound
connections, so Tailscale is the path.
- The env client is at the same commit on both machines, with the same
mujoco, gymnasium and numpy versions.
3. **The pilot's training.** Over 8,192 steps both arms ran, and the clients
exited 0.
- Per step on the server's clock: 2.94 ms local, 5.70 ms cross.
- Per call on the client's side, local then cross:
- infer wait: 2.5 ms, 4.6 ms
- env step: 0.09 ms, 0.55 ms (the laptop steps MuJoCo slower)
- feedback send: 0.07 ms, 0.45 ms
4. **Determinism.** On guangzhao, seed 0 run three times gave the same return
at every iteration. Across the two machines the returns differ from the
first iteration. The reason is the physics, not the transport:
- With no network involved, the same seed and actions give observations
that differ by 8.9e-16 after one step, 4e-7 by step 50 and order 1 by
step 100.
- This is the last bit of the floating-point math on two platforms,
amplified by a chaotic system.
- So the arms are compared by the status rule, not trajectory by
trajectory.
5. **The pilot's ladder** (one repetition, 200 exchanges).
- `lo` and `ts` are within 0.25 ms of each other at every payload.
- `phys` is 3.8 ms with states only, then 9.1, 21.8 and 55.7 ms at 48, 184
and 588 KiB.
- Pilot throughput was 19.0, 21.7 and 23.5 MB/s. With it, `phys` fits a
fixed cost plus twice the observation's bytes over the link's
throughput. The observation crosses twice per exchange: once in the
previous feedback, once in the infer.
- **This model was suggested by the pilot, and it is declared as such**.
P6 tests whether it holds on fresh measurements.

---

## Design

**Part B first**, so that the training does not share the link with it:

- **Throughput.** Five transfers of 64 MiB, laptop to guangzhao
(`throughput.sh`).
- **The ladder.** `ladder.sh`, with plugrl-protocol's conformance server on
guangzhao and `bench_client.py` on either machine:
- three rungs: `lo` and `ts` from guangzhao itself, `phys` from the laptop
- four payloads: states only, 48, 184 and 588 KiB
- five repetitions of 1,000 exchanges each; the rung order reverses on
even repetitions

**Then Part A.** The local arm and the cross arm run at the same time, seeds
0-2, with 409,600 steps each (`server_side.sh local`, `server_side.sh cross`,
`client_side.sh`):

- **The local arm**: servers and clients on guangzhao.
- **The cross arm**: servers on guangzhao, clients on the laptop.

`summarise.py` reads everything.

---

## Checks

* **V1 - the arms are what they say.**
- Every cross client connected to guangzhao's Tailscale address, and every
local client to 127.0.0.1.
- The six servers logged the same algorithm configuration apart from the
port and the name.
* **V2 - a ladder run counts** only if its client completed its 1,000
exchanges and exited 0.

---

## Predictions, and what falsifies each

**P1 - both arms run end to end**: on all six seeds, 100 iterations logged,
five checkpoints, no traceback, and the client exited 0.

**P2 - the cross arm learns**, by the status rule. **P3 - the local arm
learns**, by the same rule.

> Grounds: known item 1. E16's control learned this cell on one machine, from
> about -300 to 1,346-2,050. Where the environment is stepped changes the
> timing, not the data a step produces; item 4's differences are
> floating-point, not systematic. Each is falsified if fewer than two seeds
> clear the bar.

**P4 - the cross arm's extra time is the pilot's per-step cost.** For each
seed, the cross arm's span from iteration 1 to iteration 100 on the server's
clock minus the local arm's, against 99 x 4,096 x 2.76 ms = **1,119 s**.

> Grounds: known item 3. This fails if crossing costs something that grows
> over a run and that two iterations do not show, such as queueing or
> backlog. Holds if the difference is within **±25%** (839-1,399 s) on at
> least 2 of 3 seeds.

**P5 - `lo` and `ts` agree.** At every payload, the medians of the five
repetitions are within **0.25 ms**.

> Grounds: known item 5, and E7's `self` rung. On one machine a Tailscale
> address is short-circuited like any other local one.

**P6 - the crossing cost is fixed latency plus bytes over bandwidth.** Take
B = the median throughput of the five transfers. For 48, 184 and 588 KiB,
the median `phys` round trip is predicted as the median `phys` round trip
with states only, plus 2 x (the observation's image bytes) / B. It holds if
the measured median is within **±25%** of the prediction at all three
payloads.

> Grounds: known item 5, and suggested by it (declared above).

**Reported, not predicted:**
- every rung's pack, round-trip and unpack times
- the throughputs
- the env clients' timing breakdown in both arms
- the returns at iterations 1, 50 and 91-100 for every seed
- wall clock
- E10's ratio recomputed with `phys` at 184 KiB against the full-size
pi0.5's 100 ms forward

---

## Declared deviations allowed in advance

1. One restart of any cross-arm seed that dies for a reason outside the
experiment, such as the laptop's Wi-Fi dropping or the laptop sleeping.
It is recorded in `AMENDMENT.md`. The local arm's matching seed is rerun
with it, so that the two arms stay concurrent.

---

## Reading order

P1, V1, the status rule, P2, P3, P4; then V2, P5, P6; then the reported
figures.
Loading
Loading