Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions experiments/e39-square-fpo-plus-plus/FINDINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# E39: with the rest of FPO++'s fine-tuning, FPO stops pulling a cloned policy down on square and lifts it - short of the bar in 4.8M steps

2026-09-27/28 · Linux workstation (`guangzhao`), CPU only · one cell, three
seeds, 100 iterations of 48,000 steps · protocol: [`PROTOCOL.md`](PROTOCOL.md)
(`5c6ccd7`, after the pilot in `results/pilot.txt`, before the registered
run)

---

## The result

E35's clone of `fpo-policy`, under FPO with FPO++'s square fine-tuning in
full (#91): E37's policy loss plus FPO++'s optimizer, clipping, advantage
handling, Huber error and raw rewards. Training-rollout success:

| seed | start: iterations 1-2 | end: iterations 91-100 | gain |
| --- | --- | --- | --- |
| 0 | 0.495 | 0.787 | **+0.292** |
| 1 | 0.524 | 0.684 | +0.160 |
| 2 | 0.518 | 0.711 | +0.193 |

* **P1 holds**: 100 iterations on every seed, ten checkpoints each, no
traceback, clients exited 0. It took 7.6 hours, beside E40.
* **V1 passes**: every server ran every setting `summarise.py` lists.
* **V2 passes**: the critic learns. Explained variance over iterations
11-20 was 0.48 / 0.44 / 0.44, and 0.66-0.68 over the last ten.
* **P2 is falsified**: 1 of 3 seeds gained 0.2. The other two gained 0.16
and 0.19. By the registered rule the cell does not learn in 4.8M steps,
and `fpo-policy` · FPO · square stays "runs".

Fifty-episode evaluations after the run, with E35's harness and the value
head these checkpoints have. The harness integrates the flow's ODE and ends
an episode on success:

| | E37 (FPO++'s loss only) | E39 (FPO++'s fine-tuning) |
| --- | --- | --- |
| the clone | 0.50 | 0.50 |
| seed 0, final | 0.24 | **0.80** |
| seed 1, final | 0.30 | **0.64** |
| seed 2, final | 0.42 | **0.64** |

---

## What changed between E37 and E39

The policy fell in E37 and rose here, on every seed, in training and in
evaluation. The one thing E39 changed was FPO++'s settings beyond its loss:
- the critic's own learning rate, ten times the actor's
- AdamW
- gradient clipping
- per-minibatch advantages computed once an iteration
- raw rewards
- the Huber error
- episodes ending on success

The GAE fix (#86) came between the two runs as well. E37's critic never fit
the returns; E39's explains two thirds of them. Which of these changes did
the work is not separated: they were added together, as PROTOCOL.md says.
The fall E37 recorded was our defect, the platform running FPO without the
rest of the fine-tuning it was published with, and this removes it.

Success by window of ten iterations:

| seed | 1-10 | 11-20 | 21-30 | 31-40 | 41-50 | 51-60 | 61-70 | 71-80 | 81-90 | 91-100 |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 0 | 0.525 | 0.574 | 0.630 | 0.669 | 0.690 | 0.701 | 0.713 | 0.774 | 0.759 | 0.787 |
| 1 | 0.538 | 0.574 | 0.593 | 0.610 | 0.639 | 0.660 | 0.657 | 0.644 | 0.687 | 0.684 |
| 2 | 0.531 | 0.535 | 0.600 | 0.625 | 0.659 | 0.670 | 0.712 | 0.692 | 0.661 | 0.711 |

Every seed rose through the whole run, the last window near the best, and
none has levelled off. The clipped fraction stayed between 0.19 and 0.29
(E37: 0.42-0.50 at the end).

---

## What E39 does not show

* **That it learns by the bar.** Two seeds fell short of +0.2 at 4.8M steps.
FPO++'s own square runs are 8M steps (about 167 of these iterations).
* **Which setting mattered**, as above.
* **Anything about pi0.5** (E36), where the same FPO++ loss ran with the same
FPO defaults and collapsed.

Checkpoints and tensorboards stay on `guangzhao`. The logs, `verdicts.txt`
and `summary.tsv` are here.
157 changes: 157 additions & 0 deletions experiments/e39-square-fpo-plus-plus/PROTOCOL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
# E39 measurement protocol (pre-registered)

**Written 2026-09-27, after the pilot in `results/pilot.txt` and before the
registered run.**

This file must not be edited after the first registered data point. Anything
learned afterwards goes in `AMENDMENT.md`, dated.

---

## The question

E37 fine-tuned E35's behaviour-cloned `fpo-policy` on robomimic square
under FPO with FPO++'s policy loss. It fell: 0.50 to 0.24-0.42 over fifty
evaluation episodes. DPPO lifted the same clone to 0.78-0.90. E37's critic
never fit the returns. By this project's rule that is our defect. And E37 ran
only FPO++'s loss, not the rest of its fine-tuning: its critic learned at the
actor's rate on rewards scaled by ten, among other differences. #91 adds the
rest as options.

**Run with FPO++'s square fine-tuning in full, as far as a state-based clone
with chunks of 4 allows, does `fpo-policy` learn square under FPO?**

### What this cannot settle

* Which of the added settings mattered. They are added together.
* Comparison with FPO++'s numbers. Its policy reads images, acts in chunks of
16, and was cloned on other demonstrations.

---

## Declared in advance: what was already known

1. **E37**, same clone, FPO++'s loss only, 60 iterations. Training success
went from 0.52-0.55 to 0.26-0.39. Fifty-episode evaluations: 0.50 for the
clone, 0.24 / 0.30 / 0.42 for the three seeds' final checkpoints. Under
DPPO it went to 0.78 / 0.86 / 0.90. Value loss went from about 9 x 10⁴ to
about 4 x 10⁴, with 42-50% of samples clipped over the last ten
iterations.
2. **FPO++'s own square runs** (arXiv 2602.02481 Fig. 4, read off the plot,
approximate). About 0.28 at the start and about 0.40 at the first
evaluation. At 8M steps, about 0.58 with zero-sampling and about 0.54 with
random sampling. No FPO++ curve falls. An adapted pipeline with chunks of 4
(App. D.3) reaches about 0.9 by about 9.5M steps.
3. **The pilot** (`results/pilot.txt`): seed 0, three iterations.
- Success: 0.45, 0.55, 0.53.
- Explained variance: −0.06 (critic-only), then 0.23, 0.33.
- Clipped fraction: 21%.
- About 250 s an iteration.
4. **E37 ran with the GAE episode-boundary defect** (#86). E39 runs without
it.

---

## Design

One cell, `fpopp-square`, on `guangzhao`. Three seeds, each its own server
and three client processes of `robomimic-v1` on `square-img` (84 x 84
agentview, read by nothing). Episodes last at most 400 steps and end on
success, as FPO++'s do. The client replans every 4 steps. **100
iterations** of 12,000 chunks (48,000 steps), 4.8M steps in all.

The start is E35's clone, restored with `except-critic`, with its
observation statistics frozen (#82).

The settings are E37's policy loss plus #91's options, as `run_cell.sh`
passes them.
- Policy loss:
- chunk loss over the 4 executed steps and 7 dimensions, summed over steps
- velocity error, Huber δ = 1
- uniform flow times
- one ratio per sample, clip 0.01
- 8 samples per action, 10 flow steps
- Optimizer:
- AdamW, actor at 1e-5 with betas (0.9, 0.99)
- critic at 1e-4
- eps 1e-5, weight decay 1e-6
- each clipped to a gradient norm of 25
- Advantages:
- normalised per minibatch
- computed once per iteration (`--algo.no-fpo-playground-trick`)
- Rewards and value loss:
- raw rewards
- a truncated episode treated as ended
- value loss 0.5 x squared error
- Schedule:
- one critic-only iteration
- 10 epochs of 8 minibatches
- gamma 0.995, lambda 0.99
- Value head: two layers, 512 and 256.

Not FPO++'s:
- The value head uses SiLU, not ReLU.
- Its input is the normalised state, not image features.
- The chunk is 4, not 16.
- There are 3 client processes, not 30 environments.
- The environments are not reset at every iteration.
- The per-sample log-ratio is clamped straight-through at ±5, where FPO++'s
square run has no clamp.

Code: this branch (#91 plus this directory). Client: plugrl-env-client at
931ab56.

---

## Checks

* **V1 - the settings took**: every seed's server log restores the clone with
`except-critic` and shows every setting `summarise.py` lists. A seed that
fails is not read.
* **V2 - the critic learns** (reported; it does not decide the status): mean
explained variance over iterations 11-20 above 0.2 on every seed.

---

## The status rule

E37's. A seed's start is its mean `rollout/success` over iterations 1-2: the
clone collected them unchanged, before and during the critic-only iteration.
Its end is the mean over iterations 91-100. The cell **learns** if the end
exceeds the start by at least **+0.2** on at least **2 of 3** seeds.

---

## Predictions, and what falsifies each

**P1 - the cell runs end to end**: 100 iterations logged, at least 10
checkpoints, no traceback, clients exit 0, on every seed.

**P2 - `fpo-policy` learns square under FPO with FPO++'s square fine-tuning.**

> Grounds: known item 2. FPO++ rises on square under these settings and never
> falls. Known item 3: with them the critic starts learning within two
> iterations, where E37's did not. Against it: +0.2 from about 0.5 is a
> larger rise than FPO++'s own image-based run makes in 8M steps (about
> +0.3 from 0.28), and E39 runs 4.8M. Falsified if fewer than two seeds gain
> 0.2.

**Reported, not predicted:**
- success, explained variance and clipped fraction by window of ten
iterations
- fifty-episode evaluations of each seed's final checkpoint and of the clone,
with E35's `eval_bc.sh` as E37's were read
- wall clock

---

## Declared deviations allowed in advance

1. One restart of any run that dies for a reason outside the experiment,
recorded in `AMENDMENT.md`.

---

## Reading order

P1, V1, V2, the status rule, P2, then the reported figures.
64 changes: 64 additions & 0 deletions experiments/e39-square-fpo-plus-plus/eval50.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Fifty evaluation episodes each for E35's clone and E39's final checkpoints:
# E35's eval_bc.sh with E39's value head (512, 256) so the checkpoints load,
# and episodes that end on success. A reported figure of E39's protocol.
set -uo pipefail
P=~/zuogou/plugrl
R=$P/e39-reg/experiments/e39-square-fpo-plus-plus/results/fpopp-square/fpo/fpo-policy
O=$P/eval-e39
SERVER_PY=$P/plugrl-server/.venv/bin/python
CLIENT_DIR=$P/plugrl-env-client
CLIENT_PY=$CLIENT_DIR/.venv-robomimic/bin/python
ROOT=$P/e39-reg
mkdir -p "$O"
export OMP_NUM_THREADS=1

last() { ls "$1" | grep -E '^[0-9]+$' | sort -n | tail -1; }

run() { # NAME CKPT PORT VALUE_DIMS...
local name="$1" ckpt="$2" port="$3"; shift 3
local out="$O/$name"
mkdir -p "$out"
(cd "$ROOT" && PYTHONPATH="$ROOT/src" "$SERVER_PY" -m plugrl_server.cli \
fpo-policy default eval default \
--port "$port" --seed 0 --policy.device cpu \
--policy.obs-dim 23 --policy.action-dim 7 --policy.action-horizon 4 \
--policy.flow-steps 10 --policy.hidden-dims 1024 1024 1024 \
--policy.value-hidden-dims "$@" \
--policy.state-keys robot0_eef_pos robot0_eef_quat robot0_gripper_qpos object \
--algo.policy-checkpoint-path "$ckpt" --algo.num-episodes 50 \
--no-show-progress-bar --no-show-metric-table \
--checkpoint-base-dir "$out/ck" --exp-name "$name" --overwrite \
> "$out/server.log" 2>&1 &)
for _ in $(seq 300); do
grep -q "is listening on" "$out/server.log" 2>/dev/null && break
sleep 1
done
(cd "$CLIENT_DIR" && MUJOCO_GL=egl PYOPENGL_PLATFORM=egl "$CLIENT_PY" -m plugrl_env_client.cli robomimic-v1 \
--server-port "$port" --server-host 127.0.0.1 \
--num-envs 1 --num-episodes 50 \
--env.name square-img --env.agentview-image-size 84 84 \
--runner.replan-steps 4 --exp-name "e39eval-$name-$port" \
> "$out/client.log" 2>&1)
echo "$name client rc=$?"
pkill -f "[-]-port $port " 2>/dev/null
local summary
summary=$(ls -dt "$CLIENT_DIR"/runs/e39eval-"$name-$port"*/rollout/proc_000/summary.json 2>/dev/null | head -1)
cp "$summary" "$out/client-summary.json"
"$SERVER_PY" -c "
import json, sys
s = json.load(open(sys.argv[1]))
print(sys.argv[2], 'episodes', s['completed_episodes'], 'success rate', s['mean_success_rate'])
" "$out/client-summary.json" "$name"
}

# The clone keeps FPO's default value head, which is what it was saved with.
run clone "$P/ckpt/e35-bc/400000" 9901 256 256 256 256 256 &
sleep 5
for s in 0 1 2; do
d=$R/fpopp-square-seed$s
run "fpopp-s$s" "$d/$(last "$d")" $((9910 + s)) 512 256 &
sleep 5
done
wait
echo EVAL_E39_DONE
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"exp_name": "e39eval-clone-9901",
"process_id": 0,
"total_processes": null,
"completed_episodes": 50,
"total_episode_count": 50,
"metric_window": 100,
"mean_return": 0.5,
"mean_success_rate": 0.5,
"timing": {
"env_steps": 16144,
"infer_calls": 4045,
"feedback_calls": 4045,
"infer_wait_s": 18.98196991113946,
"infer_obs_pack_s": 0.1666160251479596,
"env_step_s": 100.17902264208533,
"feedback_s": 0.729239251697436,
"feedback_obs_pack_s": 0.23129451577551663,
"feedback_info_pack_s": 0.01254780893214047,
"collect_time_s": 120.05684783007018,
"effective_fps": 134.4696307773331,
"infer_wait_frac": 0.15810818170077856,
"infer_obs_pack_frac": 0.0013878094266125476,
"env_step_frac": 0.8344298926112057,
"feedback_frac": 0.006074116261403177
}
}
13 changes: 13 additions & 0 deletions experiments/e39-square-fpo-plus-plus/results/eval50/eval.out
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
fpopp-s0 client rc=0
Terminated
fpopp-s0 episodes 50 success rate 0.800000011920929
fpopp-s1 client rc=0
Terminated
fpopp-s1 episodes 50 success rate 0.6399999856948853
fpopp-s2 client rc=0
Terminated
fpopp-s2 episodes 50 success rate 0.6399999856948853
clone client rc=0
Terminated
clone episodes 50 success rate 0.5
EVAL_E39_DONE
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"exp_name": "e39eval-fpopp-s0-9910",
"process_id": 0,
"total_processes": null,
"completed_episodes": 50,
"total_episode_count": 50,
"metric_window": 100,
"mean_return": 0.800000011920929,
"mean_success_rate": 0.800000011920929,
"timing": {
"env_steps": 12901,
"infer_calls": 3237,
"feedback_calls": 3237,
"infer_wait_s": 15.442263034172356,
"infer_obs_pack_s": 0.13676655408926308,
"env_step_s": 79.10773805738427,
"feedback_s": 0.5646680798381567,
"feedback_obs_pack_s": 0.18654995854012668,
"feedback_info_pack_s": 0.010180247481912374,
"collect_time_s": 95.25143572548404,
"effective_fps": 135.44152801203816,
"infer_wait_frac": 0.16212105273328553,
"infer_obs_pack_frac": 0.0014358476914030586,
"env_step_frac": 0.8305149151281442,
"feedback_frac": 0.005928184447167158
}
}
Loading
Loading