Skip to content

exp: E27 - every MLP combination runs on robomimic square - #62

Merged
tactino merged 4 commits into
mainfrom
exp/e27-robomimic
Sep 27, 2026
Merged

tactino merged 4 commits into
mainfrom
exp/e27-robomimic

Conversation

@tactino

@tactino tactino commented Sep 25, 2026

Copy link
Copy Markdown
Member

The env client registers a robomimic family and dppo-policy ships a configuration for robomimic's square task, but nothing had ever run them. E27 is the robomimic column of E24's (#59) coverage matrix. Stacked on #61 (fpo-policy reads named state keys), which fpo-rm and fpodppo-rm need; the client side is PlugRL/plugrl-env-client#9.

Every MLP combination runs on robomimic square

HalfCheetah Hopper Walker2d robomimic square
fpo-policy · FPO learns (E6, E16) learns (E23) learns (E24) runs
fpo-policy · DPPO runs (E17, E18) runs (E24) runs (E24) runs
dppo-policy · DPPO runs (E24) runs (E24) runs (E24) runs
  • P1 holds, 3 of 3 cells, 9 of 9 seeds: all twenty iterations logged, checkpoint written, no traceback, client exit 0. 81,920 steps and 204 episodes per seed; 14-17 minutes per cell on 24 cores.
  • The only error in any client log is the one declared in advance: robosuite freeing its EGL context at interpreter shutdown, after the server has ended the run.

Three defects found getting here, each fixed before the registered run

  1. The robomimic extra could not be installed (robomimic 0.3.0 imports mujoco_py) - plugrl-env-client#9.
  2. No episode ever ended: 0 episodes in 4,096 steps in the first pilot. Episodes now end on success and at robomimic's 400-step horizon - plugrl-env-client#9.
  3. fpo-policy read only states["obs"] - feat(fpo-policy): read the observation from named state keys #61.

Nothing learned, and none claimed

Success 0, return 0 and 400-step episodes in every iteration of every seed. square's reward is sparse (reward_shaping: false), a random policy never succeeded once in 204 episodes, and every cell started from a random initialisation - dppo-policy can load a pretrained actor, and DPPO's own robomimic results fine-tune one, but none was given. Learning here needs a pretrained actor or a shaped reward; neither was registered.

Files

  • PROTOCOL.md (pre-registered, 8103cae, after the pilots in results/pilot.txt), FINDINGS.md
  • run.sh, run_cell.sh, summarise.py
  • summary.tsv - one row per cell and seed, in E24's columns
  • results/ - every server and client log, verdicts.txt; checkpoints and tensorboard stay on the workstation

fpo-policy read states["obs"] and nothing else, which is what the MuJoCo
client sends. The robomimic client sends one state per quantity -
robot0_eef_pos, object and the rest - so fpo-policy could not run on
robomimic at all. FPOPolicyConfig.state_keys names the keys to
concatenate, in order; the default ("obs",) is the old behaviour. A
missing key raises with the keys that were there, and widths that do not
add up to obs_dim raise instead of reaching the network.
robomimic had never run here: its extra could not be installed without
mujoco_py, and once it could, the client ran one endless episode (both
fixed in plugrl-env-client#9). E27 runs fpo-policy under FPO and DPPO
(state keys via #61) and dppo-policy under DPPO on robomimic's square
task, three seeds, twenty iterations. P1: every cell runs end to end.
P1 holds 3 of 3 cells, 9 of 9 seeds: twenty iterations logged, checkpoint
written, no traceback, client exit 0. Success 0 and 400-step episodes
throughout: square's reward is sparse and every cell started from a random
initialisation; no cell claimed to learn. Logs, verdicts and summary.tsv
copied from guangzhao; checkpoints and tensorboard stay there.
@tactino
tactino changed the base branch from feat/fpo-state-keys to main September 27, 2026 06:09
@tactino
tactino merged commit fa1d0ba into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the exp/e27-robomimic branch September 27, 2026 06:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant