Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions experiments/e27-robomimic/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
# Checkpoints and tensorboard event files; summarise.py reads out what they hold.
results/*/fpo/
results/*/dppo/
109 changes: 109 additions & 0 deletions experiments/e27-robomimic/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# E27: the platform's MLP policies run on robomimic

2026-09-25 · Linux workstation (`guangzhao`), CPU only · two policies, two
algorithms, robomimic's `square` task, three seeds per cell · protocol:
[`PROTOCOL.md`](PROTOCOL.md) · pilots: [`results/pilot.txt`](results/pilot.txt)

---

## The result

The env client registers a robomimic family and `dppo-policy` ships a
configuration and normalisation for robomimic's `square` task, but nothing
had ever run them: on the GPU cluster the client cannot even import
robomimic. E27 is the robomimic column of E24's matrix.

**All three cells run end to end, on every seed.**

| cell | policy · algorithm · task | claim | result |
| --- | --- | --- | --- |
| `fpo-rm` | `fpo-policy` · FPO · square | runs | runs, 3 of 3 (14 min) |
| `fpodppo-rm` | `fpo-policy` · DPPO · square | runs | runs, 3 of 3 (15 min) |
| `dppo-rm` | `dppo-policy` · DPPO · square | runs | runs, 3 of 3 (17 min) |

**P1 holds, 3 of 3**: on all nine seeds all twenty iterations are logged,
the checkpoint at iteration 20 is written, no server log holds a traceback
and every client exits 0. Each seed ran 81,920 environment steps as 204
episodes. The only error in any client log is the one declared in advance
(`PROTOCOL.md`, item 3): robosuite freeing its EGL context as the interpreter
shuts down, printed as "Exception ignored in", after the server has ended the
run.

With E24, the MLP policies' block of the matrix now reads:

| | HalfCheetah | Hopper | Walker2d | robomimic square |
| --- | --- | --- | --- | --- |
| `fpo-policy` · FPO | learns (E6, E16) | learns (E23) | learns (E24) | **runs (E27)** |
| `fpo-policy` · DPPO | runs (E17, E18) | runs (E24) | runs (E24) | **runs (E27)** |
| `dppo-policy` · DPPO | runs (E24) | runs (E24) | runs (E24) | **runs (E27)** |

---

## Getting it to run found three defects

Each was fixed, with tests written first, before the registered run:

1. **The `robomimic` extra could not be installed.** It pinned robomimic
0.3.0, the last on PyPI, which imports `mujoco_py` unconditionally.
Now robomimic v0.4.0 from its tag, robosuite 1.4.1 and mujoco 2.3.7
(plugrl-env-client#9).
2. **No episode ever ended.** robomimic's environments never end an episode
themselves and the client neither stopped on success nor at a horizon:
the first pilot finished 0 episodes in 4,096 steps. Episodes now end on
success and at robomimic's rollout horizon, 400 steps for `square`
(plugrl-env-client#9).
3. **`fpo-policy` could not read robomimic's observation.** It read
`states["obs"]` only; robomimic hands over named low-dimensional keys.
`--policy.state-keys` now concatenates them in order and checks the width
(#61).

---

## Nothing was learned, and nothing was expected to be

Every iteration of every seed has a success rate of 0, a return of 0 and
episodes of exactly 400 steps - the horizon. The shipped `square` metadata
sets `reward_shaping: false`, so the reward is 1 only on success. In 204
episodes per seed, a randomly initialised policy never assembled the nut
once, and a sparse reward that never fires gives an on-policy algorithm
nothing to follow.

No cell claimed to learn. `dppo-policy` can load a pretrained actor through
`checkpoint_path`, and DPPO's own robomimic results fine-tune a policy
pretrained on demonstrations; no such checkpoint was given here, so every
cell started from a random initialisation, as in E24.

---

## What this does and does not support

**Supported:**

* Both MLP policies run on robomimic's `square` task under DPPO, and
`fpo-policy` runs under FPO - every combination of these policies and
algorithms the code allows, on a fourth task family.
* robomimic now installs and runs through the client, in its own
environment, on a Linux machine with EGL.

**Not supported:**

* Anything about learning on robomimic. That needs either a pretrained
actor for `dppo-policy` or a shaped reward, and neither was registered.
* That robomimic runs on the GPU cluster: the fixed extra was installed only
on `guangzhao`.
* Seeded environments. `robomimic-v1` refuses a seed, so the three seeds vary
only the policy's initialisation.

---

## Reproducing

```bash
OMP_NUM_THREADS=1 bash run.sh # three cells, nine runs at once; about 17 minutes on 24 cores
python summarise.py # P1 cell by cell, then the figures above; writes summary.tsv
```

`results/verdicts.txt` is `summarise.py`'s output. `summary.tsv` has one row
per cell and seed in E24's columns, so a coverage figure can read both the
same way. Checkpoints and tensorboard files stay on `guangzhao`
(`.gitignore`).
111 changes: 111 additions & 0 deletions experiments/e27-robomimic/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
# E27 measurement protocol (pre-registered)

**Written 2026-09-25, after the pilot in `results/pilot.txt` and before any
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

The env client registers a robomimic family, and `dppo-policy` ships a
configuration and normalisation for robomimic's square task. Nothing had run
it. On the GPU cluster the client's environments cannot even import it:
robomimic 0.3.0, which the client's `robomimic` extra pins, imports
`mujoco_py` unconditionally, and neither cluster environment has it. E12
measured what installing it costs, not whether it runs.

**Do the platform's MLP policies train end to end on robomimic?** Three
cells, the robomimic column of E24's matrix.

### The environment, and why it is this one

On a Linux workstation (`guangzhao`), in a client venv of its own:
robomimic **v0.4.0** from its GitHub tag (PyPI stops at 0.3.0; v0.4.0 no
longer imports `mujoco_py` and supports robosuite 1.2 onwards), **robosuite
1.4.1** and **mujoco 2.3.7** - the pairing the GPU cluster already found
necessary, and the one the client's shipped environment metadata was written
for. `egl-probe`, a robomimic dependency, needs CMake to build; it was
installed into the venv, not the system. Before this file was written, the
client's own code built `NutAssemblySquare` from `square-img`, reset it and
stepped it twenty times, and its four low-dimensional keys summed to the 23
values `dppo-policy`'s square config expects.

### What this cannot settle

* Learning. Every cell claims only to run end to end.
* Seeded environments. `robomimic-v1` refuses a seed - it cannot reach the
randomness of the robosuite simulation underneath - so each cell's three
seeds vary the policy's initialisation and not the environment.

---

## Declared in advance: what was already known

1. **The first pilot** (`results/pilot.txt`): all three cells ran one
iteration and exited 0 - and finished **0 episodes in 4,096 steps**.
robomimic's environments never end an episode themselves (the shipped
metadata sets `ignore_done`); robomimic's own rollouts stop on success or
at a per-task horizon, and the client did neither. A platform defect,
fixed in plugrl-env-client#9 (931ab56) with tests written first: episodes
now end on success and at robomimic's rollout horizon, 400 for this task.
The same PR makes the `robomimic` extra installable.
2. **The second pilot**, on the fixed client: each cell ran one iteration,
finished **10 episodes** in about 4,096 steps - the 400-step horizon - and
exited 0.
3. **An error at exit that is not a failure.** Every client log ends with an
`OpenGL.error.GLError` from `eglMakeCurrent`, raised in robosuite's
`EGLGLContext.__del__` as the interpreter shuts down, after the server has
ended the run. It is printed as "Exception ignored in", changes nothing
the run did, and is not counted against P1, which reads the server's log
and the client's exit code.

---

## Design

| cell | policy | algorithm | iterations | buffer / batch | seeds |
| --- | --- | --- | --- | --- | --- |
| `fpo-rm` | `fpo-policy`, 23 state values (#61) | `fpo` | 20 | 4,096 / FPO's default | 0, 1, 2 |
| `fpodppo-rm` | `fpo-policy`, 23 state values | `dppo` default | 20 | 4,096 / 256 | 0, 1, 2 |
| `dppo-rm` | `dppo-policy`, robomimic `square` config | `dppo` default | 20 | 1,024 / 256 | 0, 1, 2 |

Task `square-img` (NutAssemblySquare), one env per client, the agentview
image at 84x84 since no policy reads images. `fpo-policy` executes every
action; `dppo-policy` its chunks of 4, so its buffer is 1,024 entries for
4,096 environment steps, as in E24. DPPO's default variant has no
learning-rate scheduler, so twenty iterations need no warmup. Code: `main`
plus #61 for the server, plugrl-env-client at 931ab56 (#9) for the client.
Every process capped at one thread, as E24 on the same machine.

---

## Predictions, and what falsifies each

**P1 - every cell runs end to end.** For each cell, on every seed: all
twenty iterations logged, the checkpoint at iteration 20, no traceback in
the server log, and a client that exits 0.

> Reported cell by cell. A cell that fails is reported as not running, with
> its error, and not rerun until it passes. Falsified, for that cell, by any
> seed failing any part.

**Reported, not predicted:** each run's first and last-ten returns, episode
lengths and success rates.

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment,
recorded in `AMENDMENT.md`.
2. A failure caused by how a cell was launched rather than what it runs is
fixed and the cell rerun once, recorded in `AMENDMENT.md`.

---

## Reading order

P1 cell by cell, then the descriptive figures.
12 changes: 12 additions & 0 deletions experiments/e27-robomimic/results/dppo-rm.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,12 @@
cell: dppo-rm = dppo-policy/default x dppo/default x robomimic square
iters: 20 x 1024 batch: 256 replan: 4 seeds: 0 1 2
server: /home/guangzhao/zuogou/plugrl/e27/../plugrl-server/.venv/bin/python (src /home/guangzhao/zuogou/plugrl/e27/src)
client: /home/guangzhao/zuogou/plugrl/e27/../plugrl-env-client/.venv-robomimic/bin/python
out: /home/guangzhao/zuogou/plugrl/e27/experiments/e27-robomimic/results/dppo-rm
start 2026-09-25 19:09:32
seed 0 finished rc=0 at 19:26:05
seed 1 finished rc=0 at 19:26:09
seed 2 finished rc=0 at 19:26:16
end 2026-09-25 19:26:16
failed seeds: 0
CELL_DONE dppo-rm
Loading
Loading