exp: E27 - every MLP combination runs on robomimic square - #62
Merged
Merged
Conversation
fpo-policy read states["obs"] and nothing else, which is what the MuJoCo
client sends. The robomimic client sends one state per quantity -
robot0_eef_pos, object and the rest - so fpo-policy could not run on
robomimic at all. FPOPolicyConfig.state_keys names the keys to
concatenate, in order; the default ("obs",) is the old behaviour. A
missing key raises with the keys that were there, and widths that do not
add up to obs_dim raise instead of reaching the network.
robomimic had never run here: its extra could not be installed without mujoco_py, and once it could, the client ran one endless episode (both fixed in plugrl-env-client#9). E27 runs fpo-policy under FPO and DPPO (state keys via #61) and dppo-policy under DPPO on robomimic's square task, three seeds, twenty iterations. P1: every cell runs end to end.
P1 holds 3 of 3 cells, 9 of 9 seeds: twenty iterations logged, checkpoint written, no traceback, client exit 0. Success 0 and 400-step episodes throughout: square's reward is sparse and every cell started from a random initialisation; no cell claimed to learn. Logs, verdicts and summary.tsv copied from guangzhao; checkpoints and tensorboard stay there.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The env client registers a robomimic family and
dppo-policyships a configuration for robomimic'ssquaretask, but nothing had ever run them. E27 is the robomimic column of E24's (#59) coverage matrix. Stacked on #61 (fpo-policyreads named state keys), whichfpo-rmandfpodppo-rmneed; the client side is PlugRL/plugrl-env-client#9.Every MLP combination runs on robomimic square
fpo-policy· FPOfpo-policy· DPPOdppo-policy· DPPOThree defects found getting here, each fixed before the registered run
robomimicextra could not be installed (robomimic 0.3.0 importsmujoco_py) - plugrl-env-client#9.fpo-policyread onlystates["obs"]- feat(fpo-policy): read the observation from named state keys #61.Nothing learned, and none claimed
Success 0, return 0 and 400-step episodes in every iteration of every seed.
square's reward is sparse (reward_shaping: false), a random policy never succeeded once in 204 episodes, and every cell started from a random initialisation -dppo-policycan load a pretrained actor, and DPPO's own robomimic results fine-tune one, but none was given. Learning here needs a pretrained actor or a shaped reward; neither was registered.Files
PROTOCOL.md(pre-registered,8103cae, after the pilots inresults/pilot.txt),FINDINGS.mdrun.sh,run_cell.sh,summarise.pysummary.tsv- one row per cell and seed, in E24's columnsresults/- every server and client log,verdicts.txt; checkpoints and tensorboard stay on the workstation