Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ on HalfCheetah with no GPU and nothing to download, is at
## What runs on it

<a href="https://plugrl.github.io/#what-runs-on-it"><img src="https://plugrl.github.io/media/coverage-grid.jpg" width="100%"
alt="Fifteen cells, four policy-algorithm pairs on HalfCheetah, Hopper, Walker2d and robomimic square, each with a frame from its trained policy, a training curve and a status. All four pairs learn the three MuJoCo tasks. On square, dppo-policy with DPPO learns, the two fpo-policy pairs run end to end, and the Gaussian policy with PPO was not run."></a>
alt="Fifteen cells, four policy-algorithm pairs on HalfCheetah, Hopper, Walker2d and robomimic square, each with a frame from its trained policy, a training curve and a status. All four pairs learn the three MuJoCo tasks. On square, both pairs trained with DPPO learn from a pretrained start, fpo-policy with FPO runs end to end but falls from its start, and the Gaussian policy with PPO was not run."></a>

Every combination of the two MLP policies and the two algorithms on four
tasks, and the baseline they are measured against: a Gaussian policy with
Expand Down
120 changes: 120 additions & 0 deletions experiments/e37-square-fpo-policy/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,120 @@
# E37: from the same behaviour-cloned start, `fpo-policy` learns square under DPPO and loses ground under our FPO

2026-09-27 · Linux workstation (`guangzhao`), CPU only · two cells, three
seeds, three client processes per seed · protocol:
[`PROTOCOL.md`](PROTOCOL.md) (`fb93a91`, after the pilots in
`results/pilot.txt`, before the registered run)

---

## The result

Both cells started from E35's clone of `fpo-policy` on DPPO's 300
demonstrations, with the demonstrations' observation statistics frozen
(#82). Training-rollout success, start against the mean of the last ten
iterations:

| cell | seed | start | end | change |
| --- | --- | --- | --- | --- |
| `dppo-square` (40 iterations) | 0 | 0.390 | 0.781 | **+0.392** |
| | 1 | 0.411 | 0.784 | **+0.372** |
| | 2 | 0.372 | 0.798 | **+0.426** |
| `fpo-square` (60 iterations) | 0 | 0.517 | 0.256 | −0.261 |
| | 1 | 0.531 | 0.353 | −0.179 |
| | 2 | 0.550 | 0.392 | −0.158 |

(The starts differ because each algorithm samples the clone its own way.
DPPO adds noise of level 1.0 at every flow step. FPO integrates the flow
from random initial noise and adds nothing along the way.)

These are training-rollout rates. So, after the run and not registered,
every final checkpoint and the clone were evaluated for fifty episodes each
with E35's `eval_bc.sh`. That script integrates the flow's ODE with no noise
along the way, ends episodes on success, and uses 10 flow steps for FPO's
checkpoints and 20 for DPPO's:

| | seed 0 | seed 1 | seed 2 |
| --- | --- | --- | --- |
| the clone, unchanged | 0.50 | | |
| `fpo-square`, final | 0.24 | 0.30 | 0.42 |
| `dppo-square`, final | 0.86 | 0.78 | 0.90 |

The fall under FPO holds under evaluation, and the rise under DPPO is larger
there than in the training rollouts. The coverage figure's clip of the
median FPO seed succeeded 4 of 5; five episodes were luck against these
fifty.

* **P1 holds**: both cells logged every iteration on every seed, with no
traceback and clients that exited 0. FPO took 5 hours and DPPO 9; both ran
together from 12:30.
* **V1 passes**: every server restored the clone with `except-critic`, ran
its cell's settings and kept the statistics frozen.
* **P3 holds**: `fpo-policy` learns square under DPPO, 3 of 3 seeds past
+0.2.
* **P2 is falsified**: under FPO with FPO++'s square fine-tuning, as far as
this FPO has it, it learns on 0 of 3 seeds. All three fell.

`fpo-policy` · DPPO · square moves from "runs" to "learns".
`fpo-policy` · FPO · square stays "runs": it runs end to end from a start
that succeeds half the time, and loses ground.

---

## The curves

Success by window of ten iterations:

| cell | seed | 1-10 | 11-20 | 21-30 | 31-40 | 41-50 | 51-60 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `dppo-square` | 0 | 0.457 | 0.600 | 0.699 | 0.781 | | |
| | 1 | 0.415 | 0.603 | 0.724 | 0.784 | | |
| | 2 | 0.472 | 0.630 | 0.724 | 0.798 | | |
| `fpo-square` | 0 | 0.518 | 0.486 | 0.424 | 0.397 | 0.318 | 0.256 |
| | 1 | 0.507 | 0.500 | 0.505 | 0.492 | 0.395 | 0.353 |
| | 2 | 0.548 | 0.543 | 0.512 | 0.443 | 0.406 | 0.392 |

DPPO was still rising at 40. FPO did not crash: it held for 20 to 40
iterations and then drifted down on every seed. E36's FPO++ arms on pi0.5
fall the same way over their first seven iterations. The first update is
harmless (E32), and what follows is not.

---

## Where the two arms differ, read from their logs (not registered)

The critics. DPPO's scales rewards by the return's running deviation and
learns at 1e-3. Its explained variance rose from about 0.40 to about 0.61
on every seed. FPO's sees rewards multiplied by 10 (FPO's `reward_scaling`)
with no normalisation, and learns at the actor's rate, 1e-5. Its value loss
fell only from about 9 x 10⁴ to about 4 x 10⁴ over sixty iterations, while
the raw advantages kept a standard deviation of about 400. FPO logs no
explained variance, so how much of the return its critic explains is not
measured here. But nothing in these logs says it learned the return, and a
critic that has not learned turns the advantages into noise. On that noise
the policy took heavily clipped steps: 43% of samples at the first update
(the pilot), 42-50% over the last ten.

PROTOCOL.md declared this among the ways the FPO arm is not FPO++'s: FPO++
gives the critic ten times the actor's learning rate, 1e-4 against 1e-5.
The same list has squared error rather than Huber, no gradient clipping,
advantages normalised over the whole buffer, no clip on the old loss, and 3
client processes rather than 30 environments.

E37 also ran on code with the GAE episode-boundary defect (#86). Here it
touches only the first and last step of each 400-step episode, since
episodes do not end on success.

---

## What E37 does not show

* **That FPO cannot fine-tune square.** By this project's rule a
platform that makes an algorithm fall is showing our defect, not a finding
about the algorithm. The critic is the first suspect, and FPO++'s own
critic settings are the first thing to try. That is a new, registered
experiment, not a reading of this one.
* **That DPPO reproduces its paper here.** These are training rollouts with
noise level 1.0, and the policy is a clone, not DPPO's diffusion policy.

Checkpoints and tensorboards stay on `guangzhao`. The logs, `verdicts.txt`
and `summary.tsv` are here.
136 changes: 136 additions & 0 deletions experiments/e37-square-fpo-policy/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# E37 measurement protocol (pre-registered)

**Written 2026-09-27, after the pilots in `results/pilot.txt` and before the
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

`fpo-policy`'s two robomimic square cells, under FPO and under DPPO, have
only ever started from random weights (E27): success 0 throughout, as it is
for any policy that starts there on square's sparse reward. Neither
algorithm's authors fine-tune square that way. DPPO fine-tunes a policy
pretrained on demonstrations; FPO++ fine-tunes behaviour-cloned flow policies.
E35 cloned `fpo-policy` on the 300 demonstrations DPPO pretrained on.

**Started from that clone, does `fpo-policy` learn square under FPO, and
under DPPO?**

### What this cannot settle

* Comparison with either paper's numbers. The success read here is that of
the training rollouts, with each algorithm's sampling; neither setting is
reproduced in full (Design).
* Whether a better clone would change the answer.

---

## Declared in advance: what was already known

1. **The clone** (E35): `fpo-policy`, chunks of 4, three hidden layers of
1024, 400,000 steps of flow matching on the demonstrations; fifty
evaluation episodes each: success 0.44 at 10 flow steps, 0.38 at 20.
2. **E34** runs `dppo-policy` from DPPO's own checkpoint in the same
environment; its start was 0.29 to 0.35 in its pilot.
3. **The first pilot** (`results/pilot.txt`): with the observation statistics
updating, FPO's critic-only iteration showed a mean ratio of 1.38 and the
next iteration's success fell from 0.51 to 0.39; #82 freezes them, and
both cells run with it.
4. **The second pilot**, statistics frozen, one seed, one client: FPO's
critic-only iteration gave a ratio of exactly 1.000 and success 0.49 then
0.50; its first update, ratio 0.993 with 43% clipped. DPPO's three
iterations by the unchanged clone succeeded 0.455, 0.295, 0.365 - swings
larger than binomial noise, as in E34 (0.27 to 0.40 between unchanged
iterations) - and its first update had `approx_kl` 0.00023 and `clipfrac`
0.10. An iteration takes about 4 minutes for FPO and 13 for DPPO with three
clients.

---

## Design

Two cells on `guangzhao`, three seeds each, run at once, each seed its own
server and three client processes of `robomimic-v1` on `square-img` (84 x 84
agentview, read by nothing), episodes of 400 steps that do not end on
success, replanning every 4 steps. Both start from E35's clone
(`--algo.restore except-critic`: the value head from the run's own
initialisation) with its observation statistics frozen.

| cell | algorithm | iterations | steps per iteration |
| --- | --- | --- | --- |
| `fpo-square` | FPO with FPO++'s square fine-tuning, as far as this FPO has it | 60 | 48,000 |
| `dppo-square` | `dppo square` (#77) with E33's changes for a flow policy | 40 | 80,000 |

`fpo-square`: FPO++'s chunk loss over the 4 executed steps and 7 dimensions,
velocity error, uniform flow times, one ratio per sample, one critic-only
iteration, clip 0.01, learning rate 1e-5, 10 epochs of 8 minibatches, 8
samples per action, gamma 0.995, lambda 0.99, 12,000 chunks per iteration, 10
flow steps. Not FPO++'s: one learning rate for actor and critic (FPO++: 1e-5
and 1e-4), squared error rather than Huber, no gradient clipping, whole-buffer
advantage normalisation, the clip of the log-ratio at 5 but none on the old
loss, and each seed's 48,000 steps collected by 3 client processes (FPO++: 30
environments).

`dppo-square`: every value of DPPO's square config (#77), with E33's noise
level 1.0 for sampling and the log-probability and 20 flow steps - the
setting under which `fpo-policy` learns all three MuJoCo tasks under DPPO -
and minibatches of 500 entries, which are DPPO's 10,000 (chunk, step)
samples at 20 steps each. Not DPPO's: all 20 steps are fine-tuned (DPPO
fine-tunes 10 of a diffusion policy's 20; `fpo-policy` has no frozen copy).

Code: this branch - #74, #77, #82 and E35's scripts merged in; client
plugrl-env-client at 931ab56.

---

## Checks

* **V1 - the settings took**: every seed's server log restores the clone
with `except-critic` and shows its cell's settings (`summarise.py` lists
them) and `freeze_obs_stats=True`.

---

## The status rule

A seed's **start** is its mean `rollout/success` over the iterations its
cloned policy collected unchanged: 1-2 for `fpo-square` (one critic-only
iteration), 1-3 for `dppo-square` (two). Its **end** is the mean over its
last ten iterations. A cell **learns** if the end exceeds the start by at
least **+0.2** on at least **2 of 3** seeds - E34's rule.

---

## Predictions, and what falsifies each

**P1 - both cells run end to end**: every iteration logged, the
checkpoints, no traceback, client exit 0, on every seed.

**P2 - `fpo-policy` learns square under FPO++'s fine-tuning.**

**P3 - `fpo-policy` learns square under DPPO's.**

> Grounds for both: each paper fine-tunes square from a cloned or
> pretrained policy of about this success and gains more than 0.2; E33 shows
> this policy class learns under DPPO in this setting. Falsified if fewer
> than two seeds gain 0.2.

**Reported, not predicted:** each seed's success per window of ten
iterations.

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment,
recorded in `AMENDMENT.md`.

---

## Reading order

P1, V1, the status rule, P2, P3, then the reported figures.
13 changes: 13 additions & 0 deletions experiments/e37-square-fpo-policy/results/dppo-square.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
cell: dppo-square = fpo-policy x dppo/square x robomimic square
iters: 40 clients per seed: 3 seeds: 0 1 2
start: /home/guangzhao/zuogou/plugrl/e37/../ckpt/e35-bc/400000 sha256 7fe7a706c2c4601d
server: /home/guangzhao/zuogou/plugrl/e37/../plugrl-server/.venv/bin/python (src /home/guangzhao/zuogou/plugrl/e37/src)
client: /home/guangzhao/zuogou/plugrl/e37/../plugrl-env-client/.venv-robomimic/bin/python
out: /home/guangzhao/zuogou/plugrl/e37/experiments/e37-square-fpo-policy/results/dppo-square
start 2026-09-27 12:31:15
seed 1 finished rc=0 at 21:15:46
seed 0 finished rc=0 at 21:20:48
seed 2 finished rc=0 at 21:23:07
end 2026-09-27 21:23:07
failed seeds: 0
CELL_DONE dppo-square
Loading
Loading