Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 83 additions & 0 deletions experiments/e41-square-fpo-plus-plus-8m/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# E41: run on to FPO++'s 8M steps, FPO with FPO++'s fine-tuning learns square by the bar - two seeds past it, the third falling back

2026-09-28 · Linux workstation (`guangzhao`), CPU only · E39's three seeds
resumed at iteration 101 and run to 167, 48,000 steps an iteration ·
protocol: [`PROTOCOL.md`](PROTOCOL.md) (`3898cce`, after the pilot in
`results/pilot.txt`, before the registered run)

---

## The result

E39's cell, each seed resumed from its own final checkpoint with
`--algo.restore all` and every setting unchanged. Training-rollout success:

| seed | start: E39's iterations 1-2 | E39's end: 91-100 | end: 158-167 | gain |
| --- | --- | --- | --- | --- |
| 0 | 0.495 | 0.787 | 0.786 | **+0.291** |
| 1 | 0.524 | 0.684 | 0.738 | **+0.214** |
| 2 | 0.518 | 0.711 | 0.672 | +0.153 |

* **P1 holds**: 67 iterations on every seed, seven checkpoints each, no
traceback, clients exited 0. It took 4.8 hours.
* **V1 passes**: every seed restored its own E39 run with `restore=all`,
logged 101 as its first iteration, and ran every setting `summarise.py`
lists.
* **The status rule**: 2 of 3 seeds gained at least 0.2, so the cell
**learns**, and `fpo-policy` · FPO · square becomes "learns".
* **P2 holds, and its grounds were wrong about which seeds.** They had
seed 2 crossing within a few iterations and seed 1 needing most of the
extension. The reverse happened: seed 1 rose 0.05 over the extension and
crossed, and seed 2 fell 0.04 below where E39 left it.

Fifty-episode evaluations of the final checkpoints, with `eval50.sh`: E39's
script with E41's runs. The clone was not run again; its 0.50 is E39's, from
the same script.

| | E39, 4.8M steps | E41, 8M steps |
| --- | --- | --- |
| the clone | 0.50 | |
| seed 0, final | 0.80 | **0.82** |
| seed 1, final | 0.64 | 0.64 |
| seed 2, final | 0.64 | 0.54 |

---

## The shape of the extension

Success by window of ten iterations, E39's last and then E41's (the last
window is seven iterations):

| seed | 91-100 | 101-110 | 111-120 | 121-130 | 131-140 | 141-150 | 151-160 | 161-167 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 0 | 0.787 | 0.809 | 0.814 | 0.817 | **0.827** | 0.809 | 0.800 | 0.778 |
| 1 | 0.684 | 0.679 | 0.709 | 0.729 | **0.760** | 0.755 | 0.728 | 0.735 |
| 2 | 0.711 | **0.744** | **0.744** | 0.738 | 0.716 | 0.707 | 0.699 | 0.658 |

E39's seeds were all still rising when it stopped. Here each rose for a
while and then came down: seed 0 by 0.05 from its best window, seed 1 by
0.03, seed 2 by 0.09. Seed 2's evaluation fell with it, from 0.64 to 0.54.
The full table from iteration 1 is in `results/verdicts.txt`.

The critic kept fitting the returns: explained variance over the last ten
iterations was 0.71 / 0.75 / 0.68, against E39's 0.66-0.68. The clipped
fraction stayed between 0.21 and 0.34, a little above E39's last ten
iterations.

---

## What E41 does not show

* **That the gain holds.** By the registered rule the cell learns, and that
is its status. But every seed came down from its best over the last thirty
to sixty iterations, and whether that is noise or the start of a decline
is not something this run can separate. As PROTOCOL.md says, there is no
further extension of this cell under this setting.
* **That a resumed run is an uninterrupted one.** The random streams
restarted at iteration 101 and the first buffer was collected fresh
(PROTOCOL.md, known item 3).
* **Which of FPO++'s settings did the work**, as in E39.
* **Anything about pi0.5.** E42 runs FPO++'s fine-tuning in full there.

Checkpoints and tensorboards stay on `guangzhao`. The logs, `verdicts.txt`,
`summary.tsv` and the evaluations are here.
107 changes: 107 additions & 0 deletions experiments/e41-square-fpo-plus-plus-8m/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# E41 measurement protocol (pre-registered)

**Written 2026-09-28, after E39's result and the pilot in
`results/pilot.txt`, before the registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question, and why it is asked now

E39 ran `fpo-policy` on square under FPO with FPO++'s square fine-tuning for
4.8M steps. Every seed rose through the whole run. One seed passed the bar,
and the other two fell short by 0.04 and 0.01. FPO++'s own square runs are
8M steps (arXiv 2602.02481, the experiments on robomimic).

**Run on to FPO++'s 8M steps, does E39's setting learn square?**

This question was decided after seeing E39's result, and it is declared as
such. What keeps it from being a search is what is fixed here and taken
from outside:
- The budget is FPO++'s published one, 8M steps: 167 iterations of 48,000,
or 8.016M.
- The start and the bar are E39's, unchanged.
- There is one extension. There will be no further one on this cell under
this setting.

---

## Declared in advance: what was already known

1. **E39**: gains of +0.292, +0.160 and +0.193 at iteration 100.
- Success over iterations 51-100 by window of ten: 0.701, 0.713, 0.774,
0.759, 0.787 / 0.660, 0.657, 0.644, 0.687, 0.684 / 0.670, 0.712,
0.692, 0.661, 0.711.
- Fifty-episode evaluations of the final checkpoints: 0.80, 0.64, 0.64,
against the clone's 0.50.
2. **The pilot** (`results/pilot.txt`): seed 0 resumed from E39's final
checkpoint with `--algo.restore all`. The server logged "Restoring from
... with restore=all", and the first iteration it logged was 101.
3. **A resume is not an uninterrupted run.** The seeds' random streams
restart. The episodes in progress when E39 stopped are not continued,
and the first iteration's buffer is collected fresh. The weights, both
AdamW groups' moments, the step count and the iteration count carry
over.

---

## Design

The E39 cell, resumed: each seed from its own E39 final checkpoint
(1200003, 1200003, 1200002) with `--algo.restore all`, run until the step
count reaches 167 x 12,000. Every setting is E39's, with the same three
client processes per seed, on `guangzhao`. Code: this branch (#91, E39,
and this directory).

---

## Checks

* **V1 - it is a resume**: every seed's server log restores its own E39
seed with `restore=all` and shows E39's settings, and its first logged
iteration is 101. A seed that fails is not read.

---

## The status rule

A seed's start is E39's: its mean `rollout/success` over E39's iterations
1-2. Its end is the mean over iterations 158-167. The cell **learns** if the
end exceeds the start by at least **+0.2** on at least **2 of 3** seeds.

---

## Predictions, and what falsifies each

**P1 - the resumed cell runs end to end**: 67 iterations logged, at least 6
checkpoints, no traceback, clients exit 0, on every seed.

**P2 - it learns within 8M steps.**

> Grounds: known item 1. Over E39's last forty iterations seed 2 rose about
> 0.04 and seed 1 about 0.02. Seed 2 is 0.007 short, and at that rate it
> passes within a few iterations. Seed 1 is 0.04 short and would need most
> of the 67, and it has been flat since iteration 60. Seed 0 is already past
> by 0.09. Two seeds past means seed 0 has to stay past and seed 2 has to
> cross. Falsified if fewer than two seeds end at least 0.2 above their
> start.

**Reported, not predicted:**
- success by window across both runs
- fifty-episode evaluations of the final checkpoints, with E39's `eval50.sh`
- wall clock

---

## Declared deviations allowed in advance

1. One restart of any seed that dies for a reason outside the experiment,
from the same E39 checkpoint, recorded in `AMENDMENT.md`.

---

## Reading order

P1, V1, the status rule, P2, then the reported figures.
62 changes: 62 additions & 0 deletions experiments/e41-square-fpo-plus-plus-8m/eval50.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
#!/usr/bin/env bash
# Fifty evaluation episodes each for E41's final checkpoints: E39's eval50.sh
# with E41's results in place of E39's. The clone is not run again; E39's
# fifty episodes of it, 0.50, used this same script. A reported figure of
# E41's protocol.
set -uo pipefail
P=~/zuogou/plugrl
R=$P/e41-reg/experiments/e41-square-fpo-plus-plus-8m/results/fpopp-square-8m/fpo/fpo-policy
O=$P/eval-e41
SERVER_PY=$P/plugrl-server/.venv/bin/python
CLIENT_DIR=$P/plugrl-env-client
CLIENT_PY=$CLIENT_DIR/.venv-robomimic/bin/python
ROOT=$P/e41-reg
mkdir -p "$O"
export OMP_NUM_THREADS=1

last() { ls "$1" | grep -E '^[0-9]+$' | sort -n | tail -1; }

run() { # NAME CKPT PORT VALUE_DIMS...
local name="$1" ckpt="$2" port="$3"; shift 3
local out="$O/$name"
mkdir -p "$out"
(cd "$ROOT" && PYTHONPATH="$ROOT/src" "$SERVER_PY" -m plugrl_server.cli \
fpo-policy default eval default \
--port "$port" --seed 0 --policy.device cpu \
--policy.obs-dim 23 --policy.action-dim 7 --policy.action-horizon 4 \
--policy.flow-steps 10 --policy.hidden-dims 1024 1024 1024 \
--policy.value-hidden-dims "$@" \
--policy.state-keys robot0_eef_pos robot0_eef_quat robot0_gripper_qpos object \
--algo.policy-checkpoint-path "$ckpt" --algo.num-episodes 50 \
--no-show-progress-bar --no-show-metric-table \
--checkpoint-base-dir "$out/ck" --exp-name "$name" --overwrite \
> "$out/server.log" 2>&1 &)
for _ in $(seq 300); do
grep -q "is listening on" "$out/server.log" 2>/dev/null && break
sleep 1
done
(cd "$CLIENT_DIR" && MUJOCO_GL=egl PYOPENGL_PLATFORM=egl "$CLIENT_PY" -m plugrl_env_client.cli robomimic-v1 \
--server-port "$port" --server-host 127.0.0.1 \
--num-envs 1 --num-episodes 50 \
--env.name square-img --env.agentview-image-size 84 84 \
--runner.replan-steps 4 --exp-name "e41eval-$name-$port" \
> "$out/client.log" 2>&1)
echo "$name client rc=$?"
pkill -f "[-]-port $port " 2>/dev/null
local summary
summary=$(ls -dt "$CLIENT_DIR"/runs/e41eval-"$name-$port"*/rollout/proc_000/summary.json 2>/dev/null | head -1)
cp "$summary" "$out/client-summary.json"
"$SERVER_PY" -c "
import json, sys
s = json.load(open(sys.argv[1]))
print(sys.argv[2], 'episodes', s['completed_episodes'], 'success rate', s['mean_success_rate'])
" "$out/client-summary.json" "$name"
}

for s in 0 1 2; do
d=$R/fpopp-square-8m-seed$s
run "fpopp8m-s$s" "$d/$(last "$d")" $((9920 + s)) 512 256 &
sleep 5
done
wait
echo EVAL_E41_DONE
10 changes: 10 additions & 0 deletions experiments/e41-square-fpo-plus-plus-8m/results/eval50/eval.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
fpopp8m-s0 client rc=0
Terminated
fpopp8m-s0 episodes 50 success rate 0.8199999928474426
fpopp8m-s1 client rc=0
Terminated
fpopp8m-s1 episodes 50 success rate 0.6399999856948853
fpopp8m-s2 client rc=0
Terminated
fpopp8m-s2 episodes 50 success rate 0.5400000214576721
EVAL_E41_DONE
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"exp_name": "e41eval-fpopp8m-s0-9920",
"process_id": 0,
"total_processes": null,
"completed_episodes": 50,
"total_episode_count": 50,
"metric_window": 100,
"mean_return": 0.8199999928474426,
"mean_success_rate": 0.8199999928474426,
"timing": {
"env_steps": 11852,
"infer_calls": 2978,
"feedback_calls": 2978,
"infer_wait_s": 13.568808950018138,
"infer_obs_pack_s": 0.12461222754791379,
"env_step_s": 71.91605789517052,
"feedback_s": 0.5407434510998428,
"feedback_obs_pack_s": 0.16805215133354068,
"feedback_info_pack_s": 0.011078746290877461,
"collect_time_s": 86.15022252383642,
"effective_fps": 137.5736434890895,
"infer_wait_frac": 0.15750172840544735,
"infer_obs_pack_frac": 0.0014464527646859593,
"env_step_frac": 0.8347750683438162,
"feedback_frac": 0.006276750486050428
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"exp_name": "e41eval-fpopp8m-s1-9921",
"process_id": 0,
"total_processes": null,
"completed_episodes": 50,
"total_episode_count": 50,
"metric_window": 100,
"mean_return": 0.6399999856948853,
"mean_success_rate": 0.6399999856948853,
"timing": {
"env_steps": 13693,
"infer_calls": 3432,
"feedback_calls": 3432,
"infer_wait_s": 15.207157797180116,
"infer_obs_pack_s": 0.14363128715194762,
"env_step_s": 83.50461922492832,
"feedback_s": 0.6178314441349357,
"feedback_obs_pack_s": 0.18975139874964952,
"feedback_info_pack_s": 0.010329514974728227,
"collect_time_s": 99.47323975339532,
"effective_fps": 137.65511240959273,
"infer_wait_frac": 0.15287687256271404,
"infer_obs_pack_frac": 0.0014439188620781304,
"env_step_frac": 0.8394681768880263,
"feedback_frac": 0.0062110316871814494
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"exp_name": "e41eval-fpopp8m-s2-9922",
"process_id": 0,
"total_processes": null,
"completed_episodes": 50,
"total_episode_count": 50,
"metric_window": 100,
"mean_return": 0.5400000214576721,
"mean_success_rate": 0.5400000214576721,
"timing": {
"env_steps": 14456,
"infer_calls": 3623,
"feedback_calls": 3623,
"infer_wait_s": 16.055397980147973,
"infer_obs_pack_s": 0.15257894713431597,
"env_step_s": 88.05673197470605,
"feedback_s": 0.6287984824739397,
"feedback_obs_pack_s": 0.20180869312025607,
"feedback_info_pack_s": 0.010526219615712762,
"collect_time_s": 104.89350738446228,
"effective_fps": 137.81596554889674,
"infer_wait_frac": 0.153063791844625,
"infer_obs_pack_frac": 0.001454608115782362,
"env_step_frac": 0.8394869632107446,
"feedback_frac": 0.005994636828848023
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
cell: fpopp-square-8m = E39's fpopp-square resumed to 167 iterations
seeds: 0 1 2 clients per seed: 3
from: /home/guangzhao/zuogou/plugrl/e41-reg/../e39-reg/experiments/e39-square-fpo-plus-plus/results/fpopp-square/fpo/fpo-policy
server: /home/guangzhao/zuogou/plugrl/e41-reg/../plugrl-server/.venv/bin/python (src /home/guangzhao/zuogou/plugrl/e41-reg/src)
client: /home/guangzhao/zuogou/plugrl/e41-reg/../plugrl-env-client/.venv-robomimic/bin/python
out: /home/guangzhao/zuogou/plugrl/e41-reg/experiments/e41-square-fpo-plus-plus-8m/results/fpopp-square-8m
start 2026-09-28 11:18:46
seed 0 resumes /home/guangzhao/zuogou/plugrl/e41-reg/../e39-reg/experiments/e39-square-fpo-plus-plus/results/fpopp-square/fpo/fpo-policy/fpopp-square-seed0/1200003
seed 1 resumes /home/guangzhao/zuogou/plugrl/e41-reg/../e39-reg/experiments/e39-square-fpo-plus-plus/results/fpopp-square/fpo/fpo-policy/fpopp-square-seed1/1200003
seed 2 resumes /home/guangzhao/zuogou/plugrl/e41-reg/../e39-reg/experiments/e39-square-fpo-plus-plus/results/fpopp-square/fpo/fpo-policy/fpopp-square-seed2/1200002
seed 0 finished rc=0 at 16:02:28
seed 1 finished rc=0 at 16:06:25
seed 2 finished rc=0 at 16:06:37
end 2026-09-28 16:06:37
failed seeds: 0
CELL_DONE fpopp-square-8m
Loading
Loading