Skip to content

figure(coverage): seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds - #80

Merged
tactino merged 4 commits into
mainfrom
figure/coverage-green
Sep 27, 2026
Merged

tactino merged 4 commits into
mainfrom
figure/coverage-green

Conversation

@tactino

@tactino tactino commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

The coverage figure's source for seven cells that changed status (the page itself is plugrl.github.io#17):

  • fpo-policy · DPPO on HalfCheetah, Hopper and Walker2d: learns (E30, E33) - with dppo-policy's structure. The evaluation commands carry --policy.action-horizon 4 --policy.flow-steps 20 --policy.hidden-dims 512 512 512 so the checkpoints load, clips replan every 4, and the training commands are E33's.
  • dppo-policy · DPPO on Hopper and Walker2d: learns (E31) at 300 iterations.
  • dppo-policy · DPPO · square: learns (E34) from DPPO's released policy; the evaluation uses dppo-policy square so the frozen copy of the early steps loads.
  • pi0-policy · FPO: holds (E32) - one update with FPO++'s chunk loss leaves pi0.5 at 33 of 50; the command is E32's, the clip that checkpoint on initial state 0.

Each clip is from the cell's median seed (E31: Hopper seed 0, Walker2d seed 2; E33: HalfCheetah 0, Hopper 3, Walker2d 0), final checkpoint, median of five evaluation episodes.

One fix to data.py: it read seeds 0-2 for every cell, and E33 ran Hopper on seeds 3-5, so that cell drew no curve. A cell can now name its seeds ("seeds": [3, 4, 5]); the rest default to 0-2 as before.

Merge after #75 (E31), #78 (E30), #79 (E33), #81 (E32) and #84 (E34): data.py refuses to write without the experiment directories a cell cites. coverage.json for the site was generated on guangzhao from this branch with those five merged in locally.

…and Walker2d learn

fpo-policy under DPPO with dppo-policy's structure (E30, E33) learns
HalfCheetah, Hopper and Walker2d; dppo-policy under DPPO learns Hopper and
Walker2d at 300 iterations (E31). Clips from each cell's median seed, the
evaluation commands carry the structure, and the training commands are the
ones those experiments ran.
E33 ran Hopper on seeds 3-5; data.py read seeds 0-2 for every cell and drew
no curve for it. Commands stay shown with seed 0.
One update leaves pi0.5 at 33 of 50 where FPO's own chunk loss took it to 0.
The clip is the E32 fpopp checkpoint on initial state 0; the command is the
one E32 trained it with.
@tactino tactino changed the title figure(coverage): the fpo-policy · DPPO row and dppo-policy's Hopper and Walker2d learn figure(coverage): fpo-policy · DPPO, dppo-policy's Hopper and Walker2d learn; pi0-policy · FPO holds Sep 27, 2026
@tactino tactino changed the title figure(coverage): fpo-policy · DPPO, dppo-policy's Hopper and Walker2d learn; pi0-policy · FPO holds figure(coverage): seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds Sep 27, 2026
@tactino
tactino merged commit ae46a17 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the figure/coverage-green branch September 27, 2026 18:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant