Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions experiments/e24-coverage/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Checkpoints and tensorboard event files: binary, and large. What they hold
# that is the result is read out by summarise.py into summary.tsv. The logs
# are committed whole.
results/*/fpo/
results/*/dppo/
results/attempt-1/*/fpo/
results/attempt-1/*/dppo/
41 changes: 41 additions & 0 deletions experiments/e24-coverage/AMENDMENT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,41 @@
# E24 amendments

Dated changes and declared deviations. `PROTOCOL.md` is not edited after the
first registered data point; anything learned afterwards goes here.

---

## 1. Moved off the Windows machine, and rerun from the start

**2026-09-25 18:15, before any cell had finished.** Declared deviation 1: a
run that dies for a reason outside the experiment is restarted once.

The first attempt started at 17:41 on the Windows machine every earlier CPU
experiment used. At 18:06 it was stopped, on the user's request, to free that
machine, which was short of memory with everything else running on it. Wave 1
had been running for 25 minutes; no seed of any cell had finished. Its logs
and checkpoints are kept, untouched and unread, in `results/attempt-1/`.

All six cells rerun from the start on a Linux workstation instead,
`guangzhao` - 24 cores, 125 GB of memory, otherwise idle - under
`~/zuogou/plugrl`, with the same `run.sh`, `run_cell.sh`, seeds, code
(`main` at 8812b54 plus this branch) and settings. What differs, and is
recorded rather than matched:

| | Windows machine | guangzhao |
| --- | --- | --- |
| OS | Windows 11 | Linux |
| torch, numpy | 2.7.1+cpu, 2.3.2 | 2.7.1+cpu, 2.3.2 |
| gymnasium, mujoco (client) | 1.2.3, 3.6.0 | 1.3.0, 3.14.0 |
| dppo | the pinned fork, 89ac416 | the pinned fork, 89ac416 |
| threads per process | torch's default | `OMP_NUM_THREADS=1` |

The thread cap is there because the machine is shared: nine server-client
pairs at torch's default would try to use every core. It changes how long an
iteration takes, not what it computes.

Nothing registered depends on the machine. P1 asks whether each combination
runs end to end, and P2's threshold, 500 against an untrained policy near 0,
is far from anything a change of simulator version could move. The pilot's
null was measured on the Windows machine; the registered runs' own first
iterations are reported beside it.
107 changes: 107 additions & 0 deletions experiments/e24-coverage/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# E24: every MuJoCo combination the code allows runs, and FPO learns Walker2d

2026-09-25 · Linux workstation (`guangzhao`), CPU only · two policies, two
algorithms, three MuJoCo tasks, three seeds per cell · protocol:
[`PROTOCOL.md`](PROTOCOL.md) · amendment: [`AMENDMENT.md`](AMENDMENT.md) -
moved off the Windows machine before any cell finished, and rerun from the start

---

## The result

PlugRL keeps the policy, the algorithm that trains it and the environment
apart. Before E24, three combinations of the two MLP policies and two
algorithms had run on MuJoCo, and DPPO's own policy class, `dppo-policy`, had
never run on this platform at all. E24 filled the rest of that block.

**All six cells run end to end, and FPO learns Walker2d on every seed.**

| cell | policy · algorithm · task | claim | result |
| --- | --- | --- | --- |
| `fpo-walker` | `fpo-policy` · FPO · Walker2d-v5 | learns | **964, 1027, 1101** from −0.3 to 0.7 |
| `fpodppo-hopper` | `fpo-policy` · DPPO · Hopper-v5 | runs | runs, 3 of 3 |
| `fpodppo-walker` | `fpo-policy` · DPPO · Walker2d-v5 | runs | runs, 3 of 3 |
| `dppo-cheetah` | `dppo-policy` · DPPO · HalfCheetah-v5 | runs | runs, 3 of 3 |
| `dppo-hopper` | `dppo-policy` · DPPO · Hopper-v5 | runs | runs, 3 of 3 |
| `dppo-walker` | `dppo-policy` · DPPO · Walker2d-v5 | runs | runs, 3 of 3 |

**P1 holds, 6 of 6**: on all eighteen seeds every iteration is logged, the
checkpoint is written, no server log holds a traceback and every client
exits 0. **P2 holds, 3 of 3**: the mean of iterations 91-100 of FPO on
Walker2d is 964.0, 1026.5 and 1101.0, against the threshold of 500 and an
untrained first iteration of −0.3 to 0.7; episodes grew from about 19 steps to
about 415.

With E6, E16 and E23 before it, the MuJoCo block now reads:

| | HalfCheetah | Hopper | Walker2d |
| --- | --- | --- | --- |
| `fpo-policy` · FPO | learns (E6, E16) | learns (E23) | **learns (E24)** |
| `fpo-policy` · DPPO | runs (E17, E18) | **runs (E24)** | **runs (E24)** |
| `dppo-policy` · DPPO | **runs (E24)** | **runs (E24)** | **runs (E24)** |
| `dppo-policy` · FPO | does not exist: FPO needs a flow policy | | |

---

## One thing the run-only cells showed, unregistered

No DPPO cell claimed to learn: E17 and E18 found DPPO did not, from a random
initialisation, in `fpo-policy`. Read as description, the twenty iterations
here say that finding belongs to `fpo-policy`, not to DPPO:

| cell | return, first → mean of last ten | episode length |
| --- | --- | --- |
| `fpo-policy` · DPPO · Hopper | 14.9, 19.7, 17.0 → 15.7, 17.6, 18.1 | ~21 → ~22 |
| `fpo-policy` · DPPO · Walker2d | 1.6, 0.0, 0.5 → 1.4, 0.1, 1.1 | ~20 → ~20 |
| `dppo-policy` · DPPO · Hopper | 11.2, 12.4, 15.8 → **75.7, 77.2, 81.5** | ~19 → **~50** |
| `dppo-policy` · DPPO · Walker2d | −2.2, −1.6, −2.6 → **220.0, 191.9, 156.5** | ~17 → **~194** |
| `dppo-policy` · DPPO · HalfCheetah | −386.3, −445.0, −382.1 → −384.0, −365.7, −418.2 | 1,000 throughout |

Under the same algorithm, the same buffer of 4,096 environment steps and the
same twenty iterations, DPPO's own policy class starts to learn Hopper and
Walker2d on every seed, and `fpo-policy` does not move. That fits E18's
untested candidate: `fpo-policy` denoises all ten steps under one small
sampling noise, while DPPO's policy clamps each step's noise from below and
predicts a chunk of four actions. It does not test it. Twenty iterations is
also far too short to call anything learned, and HalfCheetah shows nothing
either way.

---

## What this does and does not support

**Supported:**

* On this platform, both MLP policies run under DPPO on HalfCheetah, Hopper
and Walker2d, and `fpo-policy` runs under FPO on all three - every
combination of these policies, algorithms and tasks the code allows.
* FPO learns Walker2d-v5 from a random initialisation, to 964-1,101 in
409,600 steps on three seeds, with defaults tuned for nothing in
particular.
* `dppo-policy`, DPPO's own policy class with its bundled D4RL configurations
and normalisation, runs end to end here for the first time.

**Not supported:**

* That DPPO learns any of these tasks. No DPPO cell was registered to learn,
and twenty iterations cannot show it. The early rise in two
`dppo-policy` cells is a reason to run them longer, not a result.
* Anything about wall-clock cost: the machine changed mid-experiment, and
every process was capped at one thread (amendment 1).
* Anything outside the MuJoCo block. robomimic, which `dppo-policy` also has a
configuration for, does not start on the GPU cluster's environments
(`import mujoco_py` fails in both); that is recorded as a gap, not tested
here.

---

## Reproducing

```bash
OMP_NUM_THREADS=1 bash run.sh # six cells in two waves; about 21 minutes on 24 cores
python summarise.py # P1 cell by cell, P2, the figures above; writes summary.tsv
```

`summary.tsv` has one row per cell and seed and is what a coverage figure
should be drawn from. The first attempt, stopped on the Windows machine
before any cell finished, is in `results/attempt-1/` and is not read.
131 changes: 131 additions & 0 deletions experiments/e24-coverage/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,131 @@
# E24 measurement protocol (pre-registered)

**Written 2026-09-25, after the pilots in `results/pilot.txt` and before any
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

PlugRL separates the policy, the algorithm that trains it and the
environment it trains in. How many of those combinations actually run on it,
end to end? On a CPU, before today, the answer was two: `fpo-policy` with FPO
on HalfCheetah (E6, E16) and on Hopper (E23), and `fpo-policy` with DPPO on
HalfCheetah (E17-E19). DPPO's own policy class, `dppo-policy`, had never run
here at all.

E24 fills the MuJoCo block of that matrix: every combination of the two
MLP policies and the two algorithms that the code allows, on HalfCheetah,
Hopper and Walker2d. **Which of them run end to end, and does FPO also learn
Walker2d?**

### What each cell claims, and what it does not

A cell claims one of two things, never more than its evidence:

* **runs end to end** - every seed completes its iterations, writes its
checkpoint, has no traceback and a client that exits 0. Nothing about
learning. DPPO from a random initialisation did not learn HalfCheetah in
this configuration (E17, E18), so no DPPO cell here claims to learn.
* **learns** - only `fpo-policy · FPO · Walker2d`, against a threshold set
from the pilot's untrained policy.

### What the code rules out

FPO requires a flow policy (`BasePolicyGradientFlowPolicy`), and
`dppo-policy` is a diffusion policy, so **`dppo-policy · FPO` does not exist**
and is not a cell. DPPO accepts both. That leaves three policy-algorithm
pairs, and with three tasks, nine cells, three of which are already filled.

---

## Declared in advance: what was already known

1. **Filled before E24**: `fpo-policy · FPO` on HalfCheetah (learns, E6, E16)
and Hopper (learns, E23); `fpo-policy · DPPO` on HalfCheetah (runs, E17,
E18; barely changes a trained policy, E19).
2. **The pilots** (`results/pilot.txt`). `fpo-policy · FPO · Walker2d`, one
iteration on three seeds: first-iteration return **0.71, −0.28, 0.64**,
episode length about 19, all three ending cleanly. The five DPPO cells,
one seed each: the first attempt at one iteration was refused at startup
by DPPO's scheduler check, correctly - its warmup is 10 iterations - and a
second at 11 iterations was stopped once every cell had completed its
first learning updates without an error.
3. **`dppo-policy` in this repository.** Each variant reads its network, its
action chunk (4 actions) and its D4RL min-max normalisation from files
bundled under `meta/dppo`, for `halfcheetah-medium-v2`, `hopper-medium-v2`
and `walker2d-medium-v2`. The `dppo` package it builds on comes from the
pinned fork in `pyproject.toml`, installed from a local clone of that
commit; installing it added packages and changed none.

---

## Design

| cell | policy | algorithm | task | iterations | buffer / batch | seeds | claims |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `fpo-walker` | `fpo-policy` | `fpo` | Walker2d-v5 | 100 | 4,096 / FPO's default | 0, 1, 2 | learns |
| `fpodppo-hopper` | `fpo-policy` | `dppo hopper` | Hopper-v5 | 20 | 4,096 / 2,048 | 0, 1, 2 | runs |
| `fpodppo-walker` | `fpo-policy` | `dppo walker` | Walker2d-v5 | 20 | 4,096 / 2,048 | 0, 1, 2 | runs |
| `dppo-cheetah` | `dppo-policy cheetah` | `dppo cheetah` | HalfCheetah-v5 | 20 | 1,024 / 512 | 0, 1, 2 | runs |
| `dppo-hopper` | `dppo-policy hopper` | `dppo hopper` | Hopper-v5 | 20 | 1,024 / 512 | 0, 1, 2 | runs |
| `dppo-walker` | `dppo-policy walker` | `dppo walker` | Walker2d-v5 | 20 | 1,024 / 512 | 0, 1, 2 | runs |

`run_cell.sh` runs one cell; `run.sh` runs all six. Everything not in the
table is E17-E23's: CPU, one env per client, server and client seeds matched,
checkpoints every 20 iterations. `fpo-policy` executes every action it
returns (`replan-steps 1`, a chunk of 1); `dppo-policy` executes its whole
chunk of 4 before the next inference, as DPPO does, so one buffer entry is 4
environment steps. Its buffer is therefore 1,024 entries, so that an
iteration is still 4,096 environment steps, and its minibatch 512, so that
an epoch is still two minibatches, as with 4,096 and 2,048. 20 iterations
stay past the scheduler's 10-iteration warmup.

Scheduling: `fpo-walker` runs alongside `fpodppo-hopper` and
`fpodppo-walker` - nine server-client pairs - and the three `dppo-policy`
cells follow. No wall-clock prediction is registered, so contention has
nothing to invalidate.

---

## Predictions, and what falsifies each

**P1 - every cell runs end to end.** For each of the six cells, on every
seed: all its iterations logged, a checkpoint at its last save, no traceback
in the server log, and a client that exits 0.

> Reported cell by cell. A cell that fails is reported as not running, with
> its error, and the matrix shows it that way - not rerun until it passes.
> Falsified, for that cell, by any seed failing any part.

**P2 - FPO learns Walker2d.** In `fpo-walker`, the mean of iterations 91-100
is at least **500** on at least **2 of 3** seeds.

> The untrained policy scores −0.28 to 0.71 (known item 2), so 500 is far
> above anything a policy that learned nothing reaches. E23 used the same
> threshold on Hopper and every seed cleared it by more than 1,000. Falsified
> by two or more seeds below 500, and then, by the project's rule, it is a
> platform defect to locate before anything else.

**Reported, not predicted:** each cell's first and last-ten returns and
episode lengths, and each run's wall clock.

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment -
including a restart of the machine - recorded in `AMENDMENT.md`.
2. If a cell fails because of how it was launched rather than what it runs -
a wrong flag, a port in use - the launch is fixed and the cell rerun once,
recorded in `AMENDMENT.md`. A failure inside the code under test is a
result and is not rerun.

---

## Reading order

P1 cell by cell, then P2, then the descriptive figures.
4 changes: 4 additions & 0 deletions experiments/e24-coverage/results/attempt-1/fpo-walker.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
cell: fpo-walker = fpo-policy/default x fpo/default x Walker2d-v5
iters: 100 x 4096 batch: variant default replan: 1 seeds: 0 1 2
out: /d/75128/Desktop/plugrl-work/plugrl-server/.claude/worktrees/fix-pi0-image-mask-batching/experiments/e24-coverage/results/fpo-walker
start 2026-09-25 17:41:23
Loading
Loading