figure(coverage): seven cells change - fpo-policy · DPPO, dppo-policy on Hopper, Walker2d and square learn; pi0-policy · FPO holds - #80
Merged
Conversation
…and Walker2d learn fpo-policy under DPPO with dppo-policy's structure (E30, E33) learns HalfCheetah, Hopper and Walker2d; dppo-policy under DPPO learns Hopper and Walker2d at 300 iterations (E31). Clips from each cell's median seed, the evaluation commands carry the structure, and the training commands are the ones those experiments ran.
E33 ran Hopper on seeds 3-5; data.py read seeds 0-2 for every cell and drew no curve for it. Commands stay shown with seed 0.
One update leaves pi0.5 at 33 of 50 where FPO's own chunk loss took it to 0. The clip is the E32 fpopp checkpoint on initial state 0; the command is the one E32 trained it with.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The coverage figure's source for seven cells that changed status (the page itself is plugrl.github.io#17):
fpo-policy· DPPO on HalfCheetah, Hopper and Walker2d: learns (E30, E33) - withdppo-policy's structure. The evaluation commands carry--policy.action-horizon 4 --policy.flow-steps 20 --policy.hidden-dims 512 512 512so the checkpoints load, clips replan every 4, and the training commands are E33's.dppo-policy· DPPO on Hopper and Walker2d: learns (E31) at 300 iterations.dppo-policy· DPPO · square: learns (E34) from DPPO's released policy; the evaluation usesdppo-policy squareso the frozen copy of the early steps loads.pi0-policy· FPO: holds (E32) - one update with FPO++'s chunk loss leaves pi0.5 at 33 of 50; the command is E32's, the clip that checkpoint on initial state 0.Each clip is from the cell's median seed (E31: Hopper seed 0, Walker2d seed 2; E33: HalfCheetah 0, Hopper 3, Walker2d 0), final checkpoint, median of five evaluation episodes.
One fix to
data.py: it read seeds 0-2 for every cell, and E33 ran Hopper on seeds 3-5, so that cell drew no curve. A cell can now name its seeds ("seeds": [3, 4, 5]); the rest default to 0-2 as before.Merge after #75 (E31), #78 (E30), #79 (E33), #81 (E32) and #84 (E34):
data.pyrefuses to write without the experiment directories a cell cites.coverage.jsonfor the site was generated onguangzhaofrom this branch with those five merged in locally.