Skip to content

exp(e31): dppo-policy learns Hopper and Walker2d under DPPO in 300 iterations - #75

Merged
tactino merged 2 commits into
mainfrom
exp/e31-dppo-longer
Sep 27, 2026
Merged

tactino merged 2 commits into
mainfrom
exp/e31-dppo-longer

Conversation

@tactino

@tactino tactino commented Sep 27, 2026

Copy link
Copy Markdown
Member

E31, pre-registered at f8ffec0 before any run: E28's two rising dppo-policy · DPPO cells, unchanged, run to 300 iterations (1,228,800 steps per seed) against the same bar.

cell mean of iterations 291-300, seeds 0 / 1 / 2 status
dppo-policy · DPPO · Hopper 873 / 961 / 812 learns (3 of 3 past 500)
dppo-policy · DPPO · Walker2d 528 / 622 / 585 learns (3 of 3 past 500)
  • P1 (both run end to end, every seed) holds; 48 and 50 minutes per cell on guangzhao.
  • P2 (Hopper learns) holds.
  • P3 (Walker2d learns) holds. I had written that I expected it falsified, from E28's Walker2d flattening after fifty iterations. E31's first hundred iterations did not repeat E28's: at 91-100 it stood at 448 / 436 / 393 against E28's 260 / 292 / 274. The DPPO path is the same computation in both (the only changes between them are feat(dppo): make the log-probability clamp configurable #64's configurable clamp at its old default and fix(buffer): keep a per-sample scalar leaf an array when read by index #58's np.asarray), and every first iteration matches E28's exactly; runs are not bit-reproducible after the first update (E20). FINDINGS says so rather than reading a plateau into one run.

Two cells of the coverage figure change from "runs, rising" to "learns"; the figure itself is updated separately, with clips recorded from these checkpoints.

Checkpoints and tensorboards stay on guangzhao; the logs, verdicts.txt and summary.tsv are here.

E28's dppo-hopper and dppo-walker, unchanged but three times as long
(1,228,800 steps). Status rule as E28's: iterations 291-300 at least 500 on
2 of 3. P2 Hopper learns; P3 Walker2d learns, expected falsified.
…erations

Mean of iterations 291-300: Hopper 873 / 961 / 812, Walker2d 528 / 622 / 585,
3 of 3 past the bar of 500 on both. P1, P2 and P3 hold; I had expected P3
falsified from E28's flat Walker2d, which E31's first hundred iterations did
not repeat (448 / 436 / 393 at 91-100 against E28's 260 / 292 / 274, same
code path, first iterations identical).
@tactino
tactino merged commit ec93a18 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the exp/e31-dppo-longer branch September 27, 2026 18:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant