Skip to content

exp: E28 - under DPPO, dppo-policy learns HalfCheetah and climbs Hopper/Walker2d; fpo-policy does not move - #63

Merged
tactino merged 5 commits into
mainfrom
exp/e28-dppo-learns
Sep 27, 2026
Merged

tactino merged 5 commits into
mainfrom
exp/e28-dppo-learns

Conversation

@tactino

@tactino tactino commented Sep 26, 2026

Copy link
Copy Markdown
Member

E24 (#59) ran every DPPO combination on MuJoCo for twenty iterations and claimed only that each runs; unregistered, dppo-policy rose on Hopper and Walker2d while fpo-policy under the same DPPO did not move. E28 runs the same five cells, through E24's own run_cell.sh unchanged, to 100 iterations (409,600 steps), with a rule for "learns" fixed in advance. Stacked on #59.

Result

cell iterations 91-100, seeds 0 / 1 / 2 status
dppo-policy · DPPO · HalfCheetah +481 / +84 / +544 over the first learns (2 of 3 past +200)
dppo-policy · DPPO · Hopper 339 / 430 / 276 (from 11-16) did not learn in 409,600 steps (bar: 500)
dppo-policy · DPPO · Walker2d 260 / 292 / 274 (from about −2) did not learn in 409,600 steps
fpo-policy · DPPO · Hopper 25 / 25 / 27 (from 15-20) did not learn in 409,600 steps
fpo-policy · DPPO · Walker2d 3 / 1 / 2 (from 0-2) did not learn in 409,600 steps
  • P1 holds, 5 of 5 (15 seeds, 100 iterations each, no traceback, client exit 0).
  • P2 and P3 falsified: I extrapolated E24's twenty-iteration slope. Walker2d flattened after about fifty iterations; Hopper was still climbing (best iteration 493).
  • P4 holds: fpo-policy under DPPO moves by +1 to +11.
  • dppo-policy learns HalfCheetah - the one cell registered without a prediction, flat for its first twenty iterations.

The registered joint reading ("the policy class decides") needed P2 or P3 and is not claimed. Described instead: with the algorithm, variant, data per iteration and optimiser steps per iteration all equal, dppo-policy gains +84 to +544 across the three tasks and fpo-policy +1 to +20. fpo-policy learns all three under FPO, so by the project's rule its standstill under DPPO is our defect to locate. One code-read candidate, untested: dppo-policy floors every denoising step's noise at 0.1, while fpo-policy's flow steps run from 0.1 down to about 0.01 (E18's candidate, now with the code's numbers).

Every seed's first iteration matches E24's to the decimal: collection is deterministic, learning is not (E20).

Matrix after E24, E27 (#62), E28

HalfCheetah Hopper Walker2d robomimic square
fpo-policy · FPO learns learns learns runs
fpo-policy · DPPO runs, no learning runs, no learning runs, no learning runs
dppo-policy · DPPO learns runs, rising: 276-430 runs, rising: 260-292 runs

Files

  • PROTOCOL.md (pre-registered, d1f4783, before any run), FINDINGS.md
  • run.sh (calls E24's run_cell.sh), summarise.py
  • summary.tsv - one row per cell and seed, in E24's columns
  • results/ - every server and client log, verdicts.txt; checkpoints and tensorboard stay on the workstation

…s run

Fills the MuJoCo block of the platform's coverage: fpo-policy with FPO
on Walker2d (a learning claim, threshold 500 against a pilot null of
-0.28 to 0.71), fpo-policy with DPPO on Hopper and Walker2d, and
dppo-policy - never run here before - with DPPO on HalfCheetah, Hopper
and Walker2d (run-end-to-end claims). dppo-policy with FPO does not
exist: FPO needs a flow policy.
P1 holds 6 of 6: fpo-policy under FPO and DPPO, and dppo-policy - run
here for the first time - under DPPO, on HalfCheetah, Hopper and
Walker2d, eighteen seeds, no error. P2 holds 3 of 3: FPO learns
Walker2d, 964-1101 from about 0. Unregistered: in twenty iterations
dppo-policy under DPPO starts to learn Hopper and Walker2d where
fpo-policy under DPPO does not move.
E24's five DPPO cells, unchanged but run to 100 iterations. A status rule
for every cell (Hopper, Walker2d: iterations 91-100 >= 500; HalfCheetah:
+200 over the first, on 2 of 3 seeds); P2 dppo-policy learns Walker2d, P3
Hopper, P4 fpo-policy under DPPO learns neither. summarise.py smoke-tested
on E24's twenty-iteration results, whose figures it reproduces.
…licy does not move

P1 holds 5 of 5. P2 and P3 falsified: dppo-policy reaches 260-292 on
Walker2d and 276-430 on Hopper, below 500. P4 holds: fpo-policy under DPPO
ends at 25-27 and 1-3. dppo-cheetah learns by the rule (+481, +84, +544).
The registered joint reading does not apply; the contrast is described,
with the per-step noise schedule named as an untested candidate.
@tactino
tactino changed the base branch from exp/e24-coverage to main September 27, 2026 06:09
@tactino
tactino merged commit 0c5a358 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the exp/e28-dppo-learns branch September 27, 2026 06:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant