Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
93 changes: 93 additions & 0 deletions experiments/e36-pi0-longer/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,93 @@
# E36: over more iterations FPO++ runs pi0.5 into the ground, and DPPO barely moves it

2026-09-27/28 · qz103, two cards per run · pi0.5 on LIBERO-10 task 8 ·
protocol: [`PROTOCOL.md`](PROTOCOL.md) (`bb9d36b`, before any run)

---

## The result

Fifty-episode evaluations, initial states 0-49 (the untrained policy: 28-37
over seven evaluations, mean 30.9, sd 3.6):

| run | iteration 5 | iteration 10 |
| --- | --- | --- |
| `fpopp-s7` (FPO with FPO++'s chunk loss and per-sample ratio) | **5** / 50 | not run: no checkpoint |
| `fpopp-s8` | **0** / 50 | not run: no checkpoint |
| `dppo` (the `libero` variant) | 28 / 50 | **20** / 50 |

* **V1 and V2 pass.** Every run's flags took (`results/configs.txt`). Every
evaluation ran fifty episodes, its client exited 0, and it is valid
(`results/stageC_eval_correctness.tsv`).
* **P1 fails for both FPO++ runs**, and the fault is mine. The training
harness, derived from E32's, wraps the clients in E32's `timeout 43200`.
That fits two iterations. Ten FPO iterations of about 80 minutes do not
fit. At 12 hours the clients were killed and the server with them, in the
middle of the tenth learn step (`results/run.out`; the server log's
"Received SIGTERM ... the model lock was still held"). So nine iterations
learned and no iteration-10 checkpoint exists. DPPO's ten iterations of
about an hour fit, and it ran to completion.
* **P2 (FPO++ holds at iteration 10) cannot be read as registered.** Its
checkpoint does not exist. The evidence it would have read points one way:
- At iteration 5 both runs are at the collapse line or below it (5 and 0).
- The training rollouts had fallen to zero: 0.05, 0, 0 and 0, 0, 0 over
iterations 7-9.
- The one restart the protocol allows was not used. It would cost about
thirteen GPU-hours to read a number the rest already gives.
* **P3 (DPPO holds at iteration 10) holds, on the line**: 20 of 50, where
holding is 20 or more. That is also three standard deviations below the
untrained policy's mean. It went 28 at iteration 5, then 20.

Neither algorithm made pi0.5 better. The status rule's "learns" is 42 of 50
at iteration 10, and nothing came near it. `pi0-policy` · DPPO stays
"holds", now after ten iterations. `pi0-policy` · FPO held for one update
(E32) and collapsed within five.

---

## What the logs show

`results/learn-stats.txt`, per iteration:

| | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| s7 training success | 0.60 | 0.38 | 0.73 | 0.40 | 0.10 | 0.14 | 0.05 | 0 | 0 |
| s7 CFM loss of its own samples | 0.11 | 0.13 | 0.16 | 0.26 | 0.81 | 5.8 | 28 | 74 | 114 |
| s7 clipped fraction | 0 | 0.46 | 0.52 | 0.46 | 0.43 | 0.50 | 0.58 | 0.75 | 0.96 |
| s8 training success | 0.67 | 0.50 | 0.47 | 0.18 | 0.13 | 0 | 0 | 0 | 0 |
| s8 CFM loss of its own samples | 0.12 | 0.12 | 0.19 | 1.9 | 10 | 52 | 180 | 338 | 407 |

The flow-matching loss of the policy's own freshly sampled actions is the
loss the next ratio starts from. It rose a thousandfold, and the success
fell with it. The velocity field stopped describing the distribution it
samples: the policy came apart, it did not merely drift. The clipped
fraction climbing towards one says that by the end almost every sample was
outside the trust region the moment the update began.

DPPO's `approx_kl` stayed between 1.3 and 4.1 x 10⁻⁷ for all ten
iterations, and the action expert moved 0.16% by iteration 5 and 0.24% by
iteration 10 (FPO++: 2.8% by iteration 5; `results/movement.txt`). DPPO kept
pi0.5 because it hardly touched it, as in E25. Where it did end up, 20 of
50, is below where it started.

---

## Where this points

The same shape as E37 on square: FPO's first update is harmless, later ones
degrade a pretrained policy. E37's critic never fit its returns. Here the
flow model's loss on its own samples runs away. Both ran with the parts of
FPO's own defaults that FPO++ does not share (`results/configs.txt`):
- rewards scaled by 10
- values and advantages recomputed every epoch
- one learning rate for actor and critic
- no gradient clipping

E39 is running FPO++'s
square fine-tuning in full (#91) to see whether the rest of FPO++'s setting
stops the square fall. If it does, the same settings go to pi0.5 next. For
DPPO on pi0.5 the question is the other one, why it barely moves, and
fpo-policy's history (E29, E30: the noise level) is the first place to look.

Checkpoints and tensorboards stay on qz103. The results, logs and harness
outputs are here.
116 changes: 116 additions & 0 deletions experiments/e36-pi0-longer/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
# E36 measurement protocol (pre-registered)

**Written 2026-09-27, before any E36 run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

pi0.5 on LIBERO-10 task 8 now survives its first update under both
algorithms PlugRL runs on it: FPO with FPO++'s chunk loss and per-sample
ratio (E32, 33 of 50) and DPPO's `libero` variant (E25, 30 of 50 after two
iterations). Neither has been shown to make it better.

**Run ten iterations, does either improve the policy?**

### What this cannot settle

* Whether more data per update, or more updates, would. Every setting is the
one E32 and E25 ran; only the number of iterations changes.
* Other tasks. Task 8 only.

---

## Declared in advance: what was already known

1. **The untrained policy** on task 8, fifty episodes, initial states 0 to
49: 28 to 37 over seven evaluations of an unperturbed actor (E15), mean
30.9, standard deviation 3.6; 29 in E26; 36 in E32.
2. **E32's `fpopp`**: one critic-only iteration, then one update: 33 of 50;
the expert moved 1.64%.
3. **E25's DPPO**: two iterations, 30 of 50; it moved the expert's MLP and
attention about 2% as far as one FPO iteration.
4. **An iteration is 4,096 environment steps**, a few dozen episodes of up to
520 steps across ten clients; an FPO iteration takes about 78 minutes and a
DPPO one about an hour on these cards.

---

## Design

Three runs on qz103, ten iterations each, a checkpoint after every one:

| run | what | cards | server seed |
| --- | --- | --- | --- |
| `fpopp-s7` | E32's `fpopp` unchanged: FPO with FPO++'s chunk loss and per-sample ratio, the first iteration critic-only, then nine updates | 0, 1 | 7 |
| `fpopp-s8` | the same | 2, 3 | 8 |
| `dppo` | E25's cell unchanged: `pi0-policy default dppo libero`, buffer 4,096, minibatch 8 | 4, 5 | 7 |

Everything else is E32's and E25's: `pi05_libero`, ten LIBERO clients on task
8 with randomised initial states, replanning every 5 steps, client seed 7.
Then each run's iteration-5 and iteration-10 checkpoints evaluated with E32's
harness: fifty episodes, initial states 0 to 49 in order, `runner.seed` 7.
Harnesses derived from E32's `train.sh` and `eval.sh` and E25's `cell.sh` by
`derive.py` (output directories, the server seed, the code directory, and
E25's memory recorder following its cards). Code: #74 (5272832), deployed LF
as `$R/plugrl-server-e32`, which has #58.

---

## Checks

* **V1 - every run's flags took**: the config lines show E32's `fpopp`
settings for the two FPO runs and the `libero` variant with buffer 4,096
and minibatch 8 for DPPO.
* **V2 - every evaluation valid**: fifty episodes, client exit 0, `valid`
true in the harness's table.

---

## The status rule

A run's iteration-10 checkpoint **learns** at **42 of 50** or more - above
the untrained policy's mean by more than three of its standard deviations,
and five above the best it has ever scored. It **holds** at 20 or more and
**collapses** at 5 or fewer. The coverage figure's cell learns if every one
of its runs does: both `fpopp` runs for `pi0-policy` · FPO, the one DPPO run
for `pi0-policy` · DPPO.

---

## Predictions, and what falsifies each

**P1 - all three runs complete**: ten iterations, ten checkpoints, no
traceback or out-of-memory.

**P2 - FPO++ keeps pi0.5 through nine updates**: both `fpopp` runs hold at
iteration 10.

> Grounds: known item 2. Falsified if either run is at 19 or fewer.

**P3 - DPPO keeps it through ten iterations**: `dppo` holds at iteration 10.

> Grounds: known item 3 - it barely moves the policy. Falsified at 19 or
> fewer.

**Reported, not predicted:** whether any run learns - my expectation is
that none does in ten iterations of this little data; every run's success
at iterations 5 and 10; the training rollouts' success per iteration; the
movement per module group at iterations 5 and 10.

---

## Declared deviations allowed in advance

1. One restart of any run or evaluation that dies for a reason outside the
experiment, recorded in `AMENDMENT.md`.
2. Cards may be reassigned if the planned ones are occupied.

---

## Reading order

V1, V2, P1, P2, P3, the status rule, then the reported figures.
91 changes: 91 additions & 0 deletions experiments/e36-pi0-longer/derive.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
"""Derive E36's harnesses from E32's and E25's by exact substitution, and fail on a miss."""

import pathlib

HERE = pathlib.Path(__file__).resolve().parent


def derive(src: str, dst: str, subs: list[tuple[str, str]]) -> None:
text = (HERE / src).read_text(encoding="utf-8")
for old, new in subs:
n = text.count(old)
if n != 1:
raise SystemExit(f"{src}: expected exactly one of {old!r}, found {n}")
text = text.replace(old, new)
(HERE / dst).write_text(text, encoding="utf-8", newline="\n")
print(f"wrote {dst}")


derive(
"e32_train.sh",
"train.sh",
[
(
"# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]",
"# bash e36/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]\n"
"#\n"
"# E36: E32's training harness with its output under e36/ and the\n"
"# server's seed taken from SEED (default 7, E14's). E32's header follows.\n"
"#\n"
"# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]",
),
("OUT=$R/e32/$CELL", "OUT=$R/e36/$CELL"),
(" --seed 7 \\\n", " --seed ${SEED:-7} \\\n"),
('log "E32_TRAIN_DONE $CELL"', 'log "E36_TRAIN_DONE $CELL"'),
],
)

derive(
"e32_eval.sh",
"eval.sh",
[
(
"# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32",
"# E36 evaluation harness: E32's with output under e36/. E32's header:\n"
"#\n"
"# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32",
),
("E=$R/e32\n", "E=$R/e36\n"),
("RES=${E32_RES:-$E/results}", "RES=${E36_RES:-$E/results}"),
],
)

derive(
"e25_cell.sh",
"dppo_cell.sh",
[
(
"# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL.",
"# E36's DPPO cell: E25's, with output under e36/, the server on\n"
"# $R/plugrl-server-e32 (E32's code, which has #58), and the memory recorder\n"
"# following the cards the server was given. E25's header follows.\n"
"#\n"
"# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL.",
),
("E=$R/e25\n", "E=$R/e36\n"),
("SRC=$R/plugrl-server-main/src", "SRC=$R/plugrl-server-e32/src"),
(
'src=$(cat $R/plugrl-server-main/COMMIT)"',
'src=$(cat $R/plugrl-server-e32/COMMIT)"',
),
(
' cd "$R/plugrl-server-main" && exec setsid env \\',
' cd "$R/plugrl-server-e32" && exec setsid env \\',
),
(
'mkdir -p "$OUT"\n',
'mkdir -p "$OUT"\n_SG="${SRV_GPUS:-0,1}"\nexport MEM_G0="${_SG%%,*}" MEM_G1="${_SG##*,}"\n',
),
(
"--format=csv,noheader,nounits -i 0),$(nvidia-smi --query-gpu=memory.used "
"--format=csv,noheader,nounits -i 1)",
"--format=csv,noheader,nounits -i $MEM_G0),$(nvidia-smi --query-gpu=memory.used "
"--format=csv,noheader,nounits -i $MEM_G1)",
),
(
'log "peak MiB on cards 0 and 1:',
'log "peak MiB on the server cards $MEM_G0 and $MEM_G1:',
),
('log "E25_CELL_DONE $CELL"', 'log "E36_DPPO_DONE $CELL"'),
],
)
Loading
Loading