From bb9d36bc8821f1c5e774644e088ac3b2a8b2cf1b Mon Sep 17 00:00:00 2001 From: tactino <18781106300@163.com> Date: Sun, 27 Sep 2026 09:28:12 -0400 Subject: [PATCH 1/2] exp: pre-register E36 - pi0.5 under FPO++ and DPPO for ten iterations E32's fpopp on two server seeds and E25's DPPO cell, unchanged, for ten iterations; iteration-5 and iteration-10 checkpoints evaluated. Learns: 42 of 50 at iteration 10 on every run of the cell. P1 all complete, P2 FPO++ holds through nine updates, P3 DPPO holds; whether any learns is reported, expected not to. Written before any E36 run. --- experiments/e36-pi0-longer/PROTOCOL.md | 116 +++++++++++++ experiments/e36-pi0-longer/derive.py | 91 ++++++++++ experiments/e36-pi0-longer/dppo_cell.sh | 141 ++++++++++++++++ experiments/e36-pi0-longer/e25_cell.sh | 135 +++++++++++++++ experiments/e36-pi0-longer/e32_eval.sh | 213 +++++++++++++++++++++++ experiments/e36-pi0-longer/e32_train.sh | 209 +++++++++++++++++++++++ experiments/e36-pi0-longer/eval.sh | 215 ++++++++++++++++++++++++ experiments/e36-pi0-longer/movement.py | 81 +++++++++ experiments/e36-pi0-longer/run.sh | 117 +++++++++++++ experiments/e36-pi0-longer/tb_read.py | 33 ++++ experiments/e36-pi0-longer/train.sh | 214 +++++++++++++++++++++++ 11 files changed, 1565 insertions(+) create mode 100644 experiments/e36-pi0-longer/PROTOCOL.md create mode 100644 experiments/e36-pi0-longer/derive.py create mode 100644 experiments/e36-pi0-longer/dppo_cell.sh create mode 100644 experiments/e36-pi0-longer/e25_cell.sh create mode 100644 experiments/e36-pi0-longer/e32_eval.sh create mode 100644 experiments/e36-pi0-longer/e32_train.sh create mode 100644 experiments/e36-pi0-longer/eval.sh create mode 100644 experiments/e36-pi0-longer/movement.py create mode 100644 experiments/e36-pi0-longer/run.sh create mode 100644 experiments/e36-pi0-longer/tb_read.py create mode 100644 experiments/e36-pi0-longer/train.sh diff --git a/experiments/e36-pi0-longer/PROTOCOL.md b/experiments/e36-pi0-longer/PROTOCOL.md new file mode 100644 index 0000000..2e463fb --- /dev/null +++ b/experiments/e36-pi0-longer/PROTOCOL.md @@ -0,0 +1,116 @@ +# E36 measurement protocol (pre-registered) + +**Written 2026-09-27, before any E36 run.** + +This file must not be edited after the first registered data point. Anything +learned afterwards goes in `AMENDMENT.md`, dated. + +--- + +## The question + +pi0.5 on LIBERO-10 task 8 now survives its first update under both +algorithms PlugRL runs on it: FPO with FPO++'s chunk loss and per-sample +ratio (E32, 33 of 50) and DPPO's `libero` variant (E25, 30 of 50 after two +iterations). Neither has been shown to make it better. + +**Run ten iterations, does either improve the policy?** + +### What this cannot settle + +* Whether more data per update, or more updates, would. Every setting is the + one E32 and E25 ran; only the number of iterations changes. +* Other tasks. Task 8 only. + +--- + +## Declared in advance: what was already known + +1. **The untrained policy** on task 8, fifty episodes, initial states 0 to + 49: 28 to 37 over seven evaluations of an unperturbed actor (E15), mean + 30.9, standard deviation 3.6; 29 in E26; 36 in E32. +2. **E32's `fpopp`**: one critic-only iteration, then one update: 33 of 50; + the expert moved 1.64%. +3. **E25's DPPO**: two iterations, 30 of 50; it moved the expert's MLP and + attention about 2% as far as one FPO iteration. +4. **An iteration is 4,096 environment steps**, a few dozen episodes of up to + 520 steps across ten clients; an FPO iteration takes about 78 minutes and a + DPPO one about an hour on these cards. + +--- + +## Design + +Three runs on qz103, ten iterations each, a checkpoint after every one: + +| run | what | cards | server seed | +| --- | --- | --- | --- | +| `fpopp-s7` | E32's `fpopp` unchanged: FPO with FPO++'s chunk loss and per-sample ratio, the first iteration critic-only, then nine updates | 0, 1 | 7 | +| `fpopp-s8` | the same | 2, 3 | 8 | +| `dppo` | E25's cell unchanged: `pi0-policy default dppo libero`, buffer 4,096, minibatch 8 | 4, 5 | 7 | + +Everything else is E32's and E25's: `pi05_libero`, ten LIBERO clients on task +8 with randomised initial states, replanning every 5 steps, client seed 7. +Then each run's iteration-5 and iteration-10 checkpoints evaluated with E32's +harness: fifty episodes, initial states 0 to 49 in order, `runner.seed` 7. +Harnesses derived from E32's `train.sh` and `eval.sh` and E25's `cell.sh` by +`derive.py` (output directories, the server seed, the code directory, and +E25's memory recorder following its cards). Code: #74 (5272832), deployed LF +as `$R/plugrl-server-e32`, which has #58. + +--- + +## Checks + +* **V1 - every run's flags took**: the config lines show E32's `fpopp` + settings for the two FPO runs and the `libero` variant with buffer 4,096 + and minibatch 8 for DPPO. +* **V2 - every evaluation valid**: fifty episodes, client exit 0, `valid` + true in the harness's table. + +--- + +## The status rule + +A run's iteration-10 checkpoint **learns** at **42 of 50** or more - above +the untrained policy's mean by more than three of its standard deviations, +and five above the best it has ever scored. It **holds** at 20 or more and +**collapses** at 5 or fewer. The coverage figure's cell learns if every one +of its runs does: both `fpopp` runs for `pi0-policy` · FPO, the one DPPO run +for `pi0-policy` · DPPO. + +--- + +## Predictions, and what falsifies each + +**P1 - all three runs complete**: ten iterations, ten checkpoints, no +traceback or out-of-memory. + +**P2 - FPO++ keeps pi0.5 through nine updates**: both `fpopp` runs hold at +iteration 10. + +> Grounds: known item 2. Falsified if either run is at 19 or fewer. + +**P3 - DPPO keeps it through ten iterations**: `dppo` holds at iteration 10. + +> Grounds: known item 3 - it barely moves the policy. Falsified at 19 or +> fewer. + +**Reported, not predicted:** whether any run learns - my expectation is +that none does in ten iterations of this little data; every run's success +at iterations 5 and 10; the training rollouts' success per iteration; the +movement per module group at iterations 5 and 10. + +--- + +## Declared deviations allowed in advance + +1. One restart of any run or evaluation that dies for a reason outside the + experiment, recorded in `AMENDMENT.md`. +2. Cards may be reassigned if the planned ones are occupied. + +--- + +## Reading order + +V1, V2, P1, P2, P3, the status rule, then the reported figures. diff --git a/experiments/e36-pi0-longer/derive.py b/experiments/e36-pi0-longer/derive.py new file mode 100644 index 0000000..104f4b1 --- /dev/null +++ b/experiments/e36-pi0-longer/derive.py @@ -0,0 +1,91 @@ +"""Derive E36's harnesses from E32's and E25's by exact substitution, and fail on a miss.""" + +import pathlib + +HERE = pathlib.Path(__file__).resolve().parent + + +def derive(src: str, dst: str, subs: list[tuple[str, str]]) -> None: + text = (HERE / src).read_text(encoding="utf-8") + for old, new in subs: + n = text.count(old) + if n != 1: + raise SystemExit(f"{src}: expected exactly one of {old!r}, found {n}") + text = text.replace(old, new) + (HERE / dst).write_text(text, encoding="utf-8", newline="\n") + print(f"wrote {dst}") + + +derive( + "e32_train.sh", + "train.sh", + [ + ( + "# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]", + "# bash e36/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]\n" + "#\n" + "# E36: E32's training harness with its output under e36/ and the\n" + "# server's seed taken from SEED (default 7, E14's). E32's header follows.\n" + "#\n" + "# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE]", + ), + ("OUT=$R/e32/$CELL", "OUT=$R/e36/$CELL"), + (" --seed 7 \\\n", " --seed ${SEED:-7} \\\n"), + ('log "E32_TRAIN_DONE $CELL"', 'log "E36_TRAIN_DONE $CELL"'), + ], +) + +derive( + "e32_eval.sh", + "eval.sh", + [ + ( + "# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32", + "# E36 evaluation harness: E32's with output under e36/. E32's header:\n" + "#\n" + "# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32", + ), + ("E=$R/e32\n", "E=$R/e36\n"), + ("RES=${E32_RES:-$E/results}", "RES=${E36_RES:-$E/results}"), + ], +) + +derive( + "e25_cell.sh", + "dppo_cell.sh", + [ + ( + "# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL.", + "# E36's DPPO cell: E25's, with output under e36/, the server on\n" + "# $R/plugrl-server-e32 (E32's code, which has #58), and the memory recorder\n" + "# following the cards the server was given. E25's header follows.\n" + "#\n" + "# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL.", + ), + ("E=$R/e25\n", "E=$R/e36\n"), + ("SRC=$R/plugrl-server-main/src", "SRC=$R/plugrl-server-e32/src"), + ( + 'src=$(cat $R/plugrl-server-main/COMMIT)"', + 'src=$(cat $R/plugrl-server-e32/COMMIT)"', + ), + ( + ' cd "$R/plugrl-server-main" && exec setsid env \\', + ' cd "$R/plugrl-server-e32" && exec setsid env \\', + ), + ( + 'mkdir -p "$OUT"\n', + 'mkdir -p "$OUT"\n_SG="${SRV_GPUS:-0,1}"\nexport MEM_G0="${_SG%%,*}" MEM_G1="${_SG##*,}"\n', + ), + ( + "--format=csv,noheader,nounits -i 0),$(nvidia-smi --query-gpu=memory.used " + "--format=csv,noheader,nounits -i 1)", + "--format=csv,noheader,nounits -i $MEM_G0),$(nvidia-smi --query-gpu=memory.used " + "--format=csv,noheader,nounits -i $MEM_G1)", + ), + ( + 'log "peak MiB on cards 0 and 1:', + 'log "peak MiB on the server cards $MEM_G0 and $MEM_G1:', + ), + ('log "E25_CELL_DONE $CELL"', 'log "E36_DPPO_DONE $CELL"'), + ], +) diff --git a/experiments/e36-pi0-longer/dppo_cell.sh b/experiments/e36-pi0-longer/dppo_cell.sh new file mode 100644 index 0000000..0c8b56f --- /dev/null +++ b/experiments/e36-pi0-longer/dppo_cell.sh @@ -0,0 +1,141 @@ +#!/usr/bin/env bash +# E36's DPPO cell: E25's, with output under e36/, the server on +# $R/plugrl-server-e32 (E32's code, which has #58), and the memory recorder +# following the cards the server was given. E25's header follows. +# +# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL. +# +# setsid nohup bash e25/cell.sh CELL ITERS BUFFER > e25/CELL.out 2>&1 < /dev/null & +# +# E14's training harness with two changes: the algorithm is DPPO's `libero` +# variant instead of FPO, and the server runs today's `main` (8812b54, +# deployed LF to $R/plugrl-server-main) through PYTHONPATH, so the copy the +# earlier experiments ran, $R/plugrl-server, is left as it was. The env +# client is the one E14 to E22 used. Server on cards 0 and 1 (policy on the +# first), ten LIBERO clients rendering on card 2, as in E14. +# +# The minibatch is 8, E14's for FPO on this policy, not the libero variant's +# 128: each minibatch runs pi0.5's VLM prefix over three images and the +# prompt, and 128 of them do not fit on a 24 GB card (E25 pilot 2). The +# variant's 16 accumulation steps, averaged since #49, still make each +# optimizer step see 128 samples. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +E=$R/e36 +CELL="$1" +ITERS="$2" +BUFFER="$3" +OUT=$E/runs/$CELL +PORT=${PORT:-8590} +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +SRC=$R/plugrl-server-e32/src +mkdir -p "$OUT" +_SG="${SRV_GPUS:-0,1}" +export MEM_G0="${_SG%%,*}" MEM_G1="${_SG##*,}" + +log() { echo "[$(date '+%m-%d %H:%M:%S')] $*"; } +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} +kill_group() { + kill -TERM -- "-$1" 2>/dev/null + for _ in $(seq 10); do + pgrep -g "$1" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$1" 2>/dev/null +} +strays() { pgrep -f "[v]env-libero/bin/python|[p]lugrl_server.cli pi0-policy" | wc -l; } + +log "cell=$CELL iters=$ITERS buffer=$BUFFER batch=${BATCH:-8} port=$PORT src=$(cat $R/plugrl-server-e32/COMMIT)" +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif [ "$(strays)" -ne 0 ]; then + log "refusing to start: $(strays) client or server processes already running" + exit 2 +fi + +( + cd "$R/plugrl-server-e32" && exec setsid env \ + PYTHONPATH="$SRC" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" \ + CUDA_VISIBLE_DEVICES=${SRV_GPUS:-0,1} \ + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ + "$SPY" -m plugrl_server.cli pi0-policy default dppo libero \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda:0 \ + --algo.buffer-size "$BUFFER" \ + --algo.batch-size "${BATCH:-8}" \ + --algo.train-itrs "$ITERS" \ + --algo.save-interval 1 \ + --seed 7 \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +( + exec setsid bash -c 'while true; do echo "$(date +%s),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G0),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G1)"; sleep 10; done' +) > "$OUT/gpu_mem.csv" 2>/dev/null & +MPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + kill_group "$MPID" + sleep 2 + # With ALLOW_SIBLINGS the count includes other experiments' processes. + if [ "$(strays)" -eq 0 ]; then log "teardown clean"; else log "after teardown: $(strays) client or server processes on the machine"; fi +} + +for _ in $(seq 1800); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened" + tail -20 "$OUT/server.log" + teardown + exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${CLI_GPU:-2} MUJOCO_EGL_DEVICE_ID=${CLI_GPU:-2} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout 43200 "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs 10 --num-episodes 1000000000 \ + --env.task-suite-name libero_10 \ + --env.task-id 8 \ + --env.randomize-initial-state \ + --runner.replan-steps 5 --runner.seed 7 \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +# The server ends the run after ITERS learn steps; then the clients go. +while alive "$SPID"; do + alive "$CPID" || { log "clients exited while the server was running"; break; } + sleep 15 +done +log "server exited after $(( $(date +%s) - T0 ))s" +teardown + +# Learn steps are read from the checkpoints and the tensorboard by +# summarise.py; the server's log does not name them. +log "checkpoints: $(find "$OUT/ck" -name model.safetensors | sed "s|$OUT/ck/||" | sort | tr '\n' ' ')" +grep -h -E "Traceback|OutOfMemory|CUDA out of memory|Error" "$OUT/server.log" | head -5 +log "peak MiB on the server cards $MEM_G0 and $MEM_G1: $(awk -F, 'NR>0 { if ($2>a) a=$2; if ($3>b) b=$3 } END { print a", "b }' "$OUT/gpu_mem.csv")" +log "E36_DPPO_DONE $CELL" diff --git a/experiments/e36-pi0-longer/e25_cell.sh b/experiments/e36-pi0-longer/e25_cell.sh new file mode 100644 index 0000000..358775e --- /dev/null +++ b/experiments/e36-pi0-longer/e25_cell.sh @@ -0,0 +1,135 @@ +#!/usr/bin/env bash +# E25 on qz103: pi0.5 trained by DPPO on LIBERO-10 task 8, through PlugRL. +# +# setsid nohup bash e25/cell.sh CELL ITERS BUFFER > e25/CELL.out 2>&1 < /dev/null & +# +# E14's training harness with two changes: the algorithm is DPPO's `libero` +# variant instead of FPO, and the server runs today's `main` (8812b54, +# deployed LF to $R/plugrl-server-main) through PYTHONPATH, so the copy the +# earlier experiments ran, $R/plugrl-server, is left as it was. The env +# client is the one E14 to E22 used. Server on cards 0 and 1 (policy on the +# first), ten LIBERO clients rendering on card 2, as in E14. +# +# The minibatch is 8, E14's for FPO on this policy, not the libero variant's +# 128: each minibatch runs pi0.5's VLM prefix over three images and the +# prompt, and 128 of them do not fit on a 24 GB card (E25 pilot 2). The +# variant's 16 accumulation steps, averaged since #49, still make each +# optimizer step see 128 samples. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +E=$R/e25 +CELL="$1" +ITERS="$2" +BUFFER="$3" +OUT=$E/runs/$CELL +PORT=${PORT:-8590} +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +SRC=$R/plugrl-server-main/src +mkdir -p "$OUT" + +log() { echo "[$(date '+%m-%d %H:%M:%S')] $*"; } +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} +kill_group() { + kill -TERM -- "-$1" 2>/dev/null + for _ in $(seq 10); do + pgrep -g "$1" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$1" 2>/dev/null +} +strays() { pgrep -f "[v]env-libero/bin/python|[p]lugrl_server.cli pi0-policy" | wc -l; } + +log "cell=$CELL iters=$ITERS buffer=$BUFFER batch=${BATCH:-8} port=$PORT src=$(cat $R/plugrl-server-main/COMMIT)" +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif [ "$(strays)" -ne 0 ]; then + log "refusing to start: $(strays) client or server processes already running" + exit 2 +fi + +( + cd "$R/plugrl-server-main" && exec setsid env \ + PYTHONPATH="$SRC" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" \ + CUDA_VISIBLE_DEVICES=${SRV_GPUS:-0,1} \ + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ + "$SPY" -m plugrl_server.cli pi0-policy default dppo libero \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda:0 \ + --algo.buffer-size "$BUFFER" \ + --algo.batch-size "${BATCH:-8}" \ + --algo.train-itrs "$ITERS" \ + --algo.save-interval 1 \ + --seed 7 \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +( + exec setsid bash -c 'while true; do echo "$(date +%s),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i 0),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i 1)"; sleep 10; done' +) > "$OUT/gpu_mem.csv" 2>/dev/null & +MPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + kill_group "$MPID" + sleep 2 + # With ALLOW_SIBLINGS the count includes other experiments' processes. + if [ "$(strays)" -eq 0 ]; then log "teardown clean"; else log "after teardown: $(strays) client or server processes on the machine"; fi +} + +for _ in $(seq 1800); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened" + tail -20 "$OUT/server.log" + teardown + exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${CLI_GPU:-2} MUJOCO_EGL_DEVICE_ID=${CLI_GPU:-2} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout 43200 "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs 10 --num-episodes 1000000000 \ + --env.task-suite-name libero_10 \ + --env.task-id 8 \ + --env.randomize-initial-state \ + --runner.replan-steps 5 --runner.seed 7 \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +# The server ends the run after ITERS learn steps; then the clients go. +while alive "$SPID"; do + alive "$CPID" || { log "clients exited while the server was running"; break; } + sleep 15 +done +log "server exited after $(( $(date +%s) - T0 ))s" +teardown + +# Learn steps are read from the checkpoints and the tensorboard by +# summarise.py; the server's log does not name them. +log "checkpoints: $(find "$OUT/ck" -name model.safetensors | sed "s|$OUT/ck/||" | sort | tr '\n' ' ')" +grep -h -E "Traceback|OutOfMemory|CUDA out of memory|Error" "$OUT/server.log" | head -5 +log "peak MiB on cards 0 and 1: $(awk -F, 'NR>0 { if ($2>a) a=$2; if ($3>b) b=$3 } END { print a", "b }' "$OUT/gpu_mem.csv")" +log "E25_CELL_DONE $CELL" diff --git a/experiments/e36-pi0-longer/e32_eval.sh b/experiments/e36-pi0-longer/e32_eval.sh new file mode 100644 index 0000000..648495c --- /dev/null +++ b/experiments/e36-pi0-longer/e32_eval.sh @@ -0,0 +1,213 @@ +#!/usr/bin/env bash +# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32 +# through PYTHONPATH, output under e32/, and the source manifest taken +# through the same PYTHONPATH - E25's and E26's hashed the older copy the +# editable install points at. E14's own header follows. +# +# E14 evaluation harness. Derived by sed from e11/e11_stageC_eval.sh; only the +# output paths differ, so the measurement is E11's, unchanged. +# +# bash e11_stageC_eval.sh CELL POLICY_LABEL TASK_ID EPISODES PORT [BUDGET_SECONDS] [CHECKPOINT_DIR] +# +# Evaluates one policy on one libero_10 task: the baseline when CHECKPOINT_DIR +# is empty, a fine-tuned checkpoint otherwise. One env client process runs +# EPISODES episodes with initial states taken in order, 0 to EPISODES-1, so the +# baseline and the fine-tuned policy face exactly the same initial states. This +# is also how openpi evaluates LIBERO. Ten processes on one task would each +# start again from initial state 0 and repeat the same few states. +# +# Process hygiene is the Stage A harness's. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +E=$R/e32 +RES=${E32_RES:-$E/results} +mkdir -p "$RES" + +CELL="$1" +POLICY_LABEL="$2" +TASK_ID="$3" +NEP="$4" +PORT="$5" +BUDGET_S="${6:-28800}" +CKPT_DIR="${7:-}" +SUITE=libero_10 +NPROC=1 + +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +UV=$HOME/.local/bin/uv +OUT=$E/runs-$CELL + +CLIENT_PATTERN="[v]env-libero/bin/python" +SERVER_PATTERN="[p]lugrl_server.cli pi0-policy" + +log() { echo "[$(date +%H:%M:%S)] $*"; } + +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} + +free_port() { + local p="$1" + for pid in $(ss -tlnp 2>/dev/null | awk -v pat=":$p" '$4 ~ pat {print $NF}' | grep -oP 'pid=\K[0-9]+' | sort -u); do + kill -9 "$pid" 2>/dev/null + done + sleep 1 +} + +kill_group() { + local pgid="$1" + kill -TERM -- "-$pgid" 2>/dev/null + for _ in $(seq 10); do + pgrep -g "$pgid" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$pgid" 2>/dev/null + sleep 1 +} + +report_strays() { + local what="$1" pattern="$2" pids + pids=$(pgrep -d, -f "$pattern") + if [ -n "$pids" ]; then + log "$what: $(echo "$pids" | tr ',' '\n' | wc -l) process(es) alive" + ps -o pid,ppid,pgid,etimes,args -p "$pids" | cut -c1-180 + return 1 + fi + return 0 +} + +source_manifest() { + local py="$1" pkg="$2" dir + dir=$("$py" -c "import importlib.util as u; print(u.find_spec('$pkg').submodule_search_locations[0])" 2>/dev/null | tail -1) + if [ -z "$dir" ] || [ ! -d "$dir" ]; then + echo "unavailable $pkg" + return + fi + (cd "$dir" && find . -name '*.py' -not -path '*/__pycache__/*' | sort | while read -r f; do + printf '%s %s/%s\n' "$(tr -d '\r' < "$f" | sha256sum | cut -d' ' -f1)" "$pkg" "${f#./}" + done) +} + +log "cell=$CELL policy=$POLICY_LABEL suite=$SUITE task_id=$TASK_ID episodes=$NEP procs=$NPROC port=$PORT budget=${BUDGET_S}s checkpoint=${CKPT_DIR:-none} res=$RES" + +if [ -n "$CKPT_DIR" ] && [ ! -f "$CKPT_DIR/model.safetensors" ]; then + log "no model.safetensors in $CKPT_DIR" + exit 2 +fi + +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif ! report_strays "refusing to start, env client processes" "$CLIENT_PATTERN" \ + || ! report_strays "refusing to start, pi0 server processes" "$SERVER_PATTERN"; then + exit 2 +fi + +rm -rf "$OUT" +mkdir -p "$OUT" + +"$UV" pip freeze --python "$SPY" > "$RES/environment-server.txt" 2>/dev/null +"$UV" pip freeze --python "$CPY" > "$RES/environment-client.txt" 2>/dev/null +nvidia-smi --query-gpu=index,name,memory.total,driver_version --format=csv > "$RES/environment-gpu.txt" 2>/dev/null +{ + PYTHONPATH="$R/plugrl-server-e32/src" source_manifest "$SPY" plugrl_server + source_manifest "$CPY" plugrl_env_client +} > "$RES/environment-source-$CELL.txt" +log "source manifest: $(wc -l < "$RES/environment-source-$CELL.txt") files, sha256 $(sha256sum "$RES/environment-source-$CELL.txt" | cut -c1-16)" +if [ -n "$CKPT_DIR" ]; then + log "checkpoint sha256: $(sha256sum "$CKPT_DIR/model.safetensors" | cut -c1-16)" +fi + +CKPT_ARGS=() +if [ -n "$CKPT_DIR" ]; then + CKPT_ARGS=(--algo.policy-checkpoint-path "$CKPT_DIR") +fi + +free_port "$PORT" + +( + cd "$R/plugrl-server-e32" && exec setsid env \ + PYTHONPATH="$R/plugrl-server-e32/src" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" CUDA_VISIBLE_DEVICES=${EVAL_SRV_GPU:-0} \ + "$SPY" -m plugrl_server.cli pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda \ + "${CKPT_ARGS[@]}" \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + free_port "$PORT" + if report_strays "after teardown, env client processes" "$CLIENT_PATTERN" \ + && report_strays "after teardown, pi0 server processes" "$SERVER_PATTERN"; then + log "teardown clean" + else + log "TEARDOWN_NOT_CLEAN" + fi +} + +# Cold, after hours out of Lustre's cache, the server has taken over 600 s to +# import and construct. +for _ in $(seq 1500); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened" + tail -20 "$OUT/server.log" + teardown + exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${EVAL_CLI_GPU:-1} MUJOCO_EGL_DEVICE_ID=${EVAL_CLI_GPU:-1} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout "$BUDGET_S" "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs "$NPROC" --num-episodes "$NEP" \ + --env.task-suite-name "$SUITE" \ + --env.task-id "$TASK_ID" \ + --env.no-randomize-initial-state \ + --runner.replan-steps 5 \ + --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +SERVER_DIED=0 +while alive "$CPID"; do + if ! alive "$SPID"; then + log "server exited while the client was running" + SERVER_DIED=1 + kill_group "$CPID" + break + fi + sleep 5 +done +wait "$CPID" +CEXIT=$? +[ "$SERVER_DIED" = 1 ] && [ "$CEXIT" = 0 ] && CEXIT=97 +WALL=$(( $(date +%s) - T0 )) +log "client exited with $CEXIT after ${WALL}s" + +teardown + +"$SPY" "$R/e11/e11_stageC_eval_post.py" --cell "$CELL" --policy-label "$POLICY_LABEL" --task-id "$TASK_ID" \ + --nep "$NEP" --out "$OUT" --res "$RES" --client-exit "$CEXIT" --wall "$WALL" +log "STAGEC_EVAL_DONE $CELL" diff --git a/experiments/e36-pi0-longer/e32_train.sh b/experiments/e36-pi0-longer/e32_train.sh new file mode 100644 index 0000000..bbe6998 --- /dev/null +++ b/experiments/e36-pi0-longer/e32_train.sh @@ -0,0 +1,209 @@ +#!/usr/bin/env bash +# Does a second learn step fit once the optimizer's state is on another card, +# and what does the transfer cost? +# +# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE] +# +# E32: E14's training harness, unchanged but for three things - the +# server runs $R/plugrl-server-e32 (main plus #74, deployed LF) through +# PYTHONPATH, output goes under e32/, and ARM_FLAGS, when set, is split +# on spaces and passed to the server after every other flag. +# +# E11 stopped at one iteration of ten: the float32 copies and Adam's moments +# do not exist until the first learn step and then have to fit beside +# everything the second one needs. Measured, that state is 3,578 MiB; moving +# it to a second card leaves the model's card with 8,046 MiB instead of +# 11,624. +# +# Two iterations is enough to answer the question, since it was the second +# that died. Ten is the experiment, and only worth starting if this passes. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +CELL="$1" +ITERATIONS="${2:-2}" +PORT="${3:-8161}" +MASTER_DEVICE="${4:-cuda:1}" + +BUFFER_SIZE=${BUFFER_SIZE:-4096} +BATCH_SIZE=8 +N_SAMPLES=4 +TASK_ID=8 +SUITE=libero_10 +NPROC=10 +GLOBAL_STEPS=$(( ITERATIONS * BUFFER_SIZE )) + +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +OUT=$R/e32/$CELL + +CLIENT_PATTERN="[v]env-libero.*/bin/python" +SERVER_PATTERN="[p]lugrl_server.cli pi0-policy" + +# The memory recorder samples by physical index, which CUDA_VISIBLE_DEVICES +# does not remap. Follow the cards the run was actually given. +_SG="${SRV_GPUS:-0,1}" +export MEM_G0="${_SG%%,*}" +export MEM_G1="${_SG##*,}" + +log() { echo "[$(date +%H:%M:%S)] $*"; } + +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} + +kill_group() { + local pgid="$1" + kill -TERM -- "-$pgid" 2>/dev/null + for _ in $(seq 30); do + pgrep -g "$pgid" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$pgid" 2>/dev/null + sleep 1 +} + +report_strays() { + local what="$1" pattern="$2" pids + pids=$(pgrep -f "$pattern" | while read -r p; do + [ "$(ps -o comm= -p "$p" 2>/dev/null)" = python ] && echo "$p" + done | paste -sd, -) + if [ -n "$pids" ]; then + log "$what: $(echo "$pids" | tr ',' '\n' | wc -l) alive" + ps -o pid,etimes,args -p "$pids" | cut -c1-120 | head -5 + return 1 + fi + return 0 +} + +log "cell=$CELL iterations=$ITERATIONS master_device=$MASTER_DEVICE lr=${LR:-1e-5} updates=${UPDATES:-4} clip=${CLIP_EPS:-0.05} drift=${DRIFT:-0} warmup=${WARMUP:-0} srv_gpus=${SRV_GPUS:-0,1} cli_gpu=${CLI_GPU:-2} buffer=$BUFFER_SIZE batch=$BATCH_SIZE" + +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif ! report_strays "refusing to start, clients" "$CLIENT_PATTERN" \ + || ! report_strays "refusing to start, servers" "$SERVER_PATTERN"; then + exit 2 +fi + +rm -rf "$OUT"; mkdir -p "$OUT" + +MASTER_FLAG=() +[ "$MASTER_DEVICE" != "none" ] && MASTER_FLAG=(--algo.master-weights-device "$MASTER_DEVICE") + +# Only passed when set, so the harness still runs against a server that does +# not have the flag at all. +OPT_FLAGS=() +[ -n "${WARMUP:-}" ] && [ "${WARMUP:-0}" != 0 ] \ + && OPT_FLAGS+=(--algo.critic-warmup-iterations "$WARMUP") +[ -n "${DRIFT:-}" ] && [ "${DRIFT:-0}" != 0 ] \ + && OPT_FLAGS+=(--algo.max-policy-drift "$DRIFT") +if [ -n "${ARM_FLAGS:-}" ]; then + read -r -a _ARM <<< "$ARM_FLAGS" + OPT_FLAGS+=("${_ARM[@]}") +fi +log "arm flags: ${ARM_FLAGS:-none}" + +( + cd "$R/plugrl-server-e32" && exec setsid env \ + PYTHONPATH="$R/plugrl-server-e32/src" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" \ + CUDA_VISIBLE_DEVICES=${SRV_GPUS:-0,1} \ + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ + "$SPY" -m plugrl_server.cli pi0-policy default fpo default \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda:0 \ + --algo.learning-rate ${LR:-1e-5} \ + --algo.batch-size "$BATCH_SIZE" \ + --algo.num-updates-per-batch ${UPDATES:-4} \ + --algo.n-samples-per-action "$N_SAMPLES" \ + --algo.buffer-size "$BUFFER_SIZE" \ + --algo.clipping-epsilon ${CLIP_EPS:-0.05} \ + --algo.global-steps "$GLOBAL_STEPS" \ + --algo.save-interval 1 \ + "${MASTER_FLAG[@]}" \ + "${OPT_FLAGS[@]}" \ + --seed 7 \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +# Peak memory on both cards, sampled from outside: nothing in the server +# records it, and which card holds what is the point here. +( + exec setsid bash -c 'while true; do echo "$(date +%s),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G0),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G1)"; sleep 10; done' +) > "$OUT/gpu_mem.csv" 2>/dev/null & +MPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + kill_group "$MPID" + if report_strays "after teardown, clients" "$CLIENT_PATTERN" \ + && report_strays "after teardown, servers" "$SERVER_PATTERN"; then + log "teardown clean" + else + log "TEARDOWN_NOT_CLEAN" + fi +} + +for _ in $(seq 1800); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened"; tail -15 "$OUT/server.log"; teardown; exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${CLI_GPU:-2} MUJOCO_EGL_DEVICE_ID=${CLI_GPU:-2} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout 43200 "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs "$NPROC" --num-episodes 1000000000 \ + --env.task-suite-name "$SUITE" \ + --env.task-id "$TASK_ID" \ + --env.randomize-initial-state \ + --runner.replan-steps 5 --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +SERVER_DIED=0 +while alive "$CPID"; do + if ! alive "$SPID"; then + log "server exited" + SERVER_DIED=1 + for _ in $(seq 60); do alive "$CPID" || break; sleep 1; done + kill_group "$CPID" + break + fi + sleep 15 +done +wait "$CPID" +CEXIT=$? +WALL=$(( $(date +%s) - T0 )) +log "clients exit $CEXIT after ${WALL}s (server_died=$SERVER_DIED)" + +teardown + +log "=== learn steps seen ===" +grep -cE "learn|Learn" "$OUT/server.log" 2>/dev/null | head -1 +log "=== out of memory? ===" +grep -c "OutOfMemoryError" "$OUT/server.log" 2>/dev/null | head -1 +log "=== peak memory, GPU0 and GPU1 ===" +awk -F, 'NR>0 {if ($2>m0) m0=$2; if ($3>m1) m1=$3} END {print "gpu0_peak_mib=" m0 " gpu1_peak_mib=" m1}' "$OUT/gpu_mem.csv" 2>/dev/null +log "=== checkpoints ===" +find "$OUT/ck" -name model.safetensors -printf '%TH:%TM %s %p\n' 2>/dev/null | sort | tail -5 +log "E32_TRAIN_DONE $CELL" diff --git a/experiments/e36-pi0-longer/eval.sh b/experiments/e36-pi0-longer/eval.sh new file mode 100644 index 0000000..da1cff8 --- /dev/null +++ b/experiments/e36-pi0-longer/eval.sh @@ -0,0 +1,215 @@ +#!/usr/bin/env bash +# E36 evaluation harness: E32's with output under e36/. E32's header: +# +# E32 evaluation harness: E14's, with the server on $R/plugrl-server-e32 +# through PYTHONPATH, output under e32/, and the source manifest taken +# through the same PYTHONPATH - E25's and E26's hashed the older copy the +# editable install points at. E14's own header follows. +# +# E14 evaluation harness. Derived by sed from e11/e11_stageC_eval.sh; only the +# output paths differ, so the measurement is E11's, unchanged. +# +# bash e11_stageC_eval.sh CELL POLICY_LABEL TASK_ID EPISODES PORT [BUDGET_SECONDS] [CHECKPOINT_DIR] +# +# Evaluates one policy on one libero_10 task: the baseline when CHECKPOINT_DIR +# is empty, a fine-tuned checkpoint otherwise. One env client process runs +# EPISODES episodes with initial states taken in order, 0 to EPISODES-1, so the +# baseline and the fine-tuned policy face exactly the same initial states. This +# is also how openpi evaluates LIBERO. Ten processes on one task would each +# start again from initial state 0 and repeat the same few states. +# +# Process hygiene is the Stage A harness's. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +E=$R/e36 +RES=${E36_RES:-$E/results} +mkdir -p "$RES" + +CELL="$1" +POLICY_LABEL="$2" +TASK_ID="$3" +NEP="$4" +PORT="$5" +BUDGET_S="${6:-28800}" +CKPT_DIR="${7:-}" +SUITE=libero_10 +NPROC=1 + +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +UV=$HOME/.local/bin/uv +OUT=$E/runs-$CELL + +CLIENT_PATTERN="[v]env-libero/bin/python" +SERVER_PATTERN="[p]lugrl_server.cli pi0-policy" + +log() { echo "[$(date +%H:%M:%S)] $*"; } + +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} + +free_port() { + local p="$1" + for pid in $(ss -tlnp 2>/dev/null | awk -v pat=":$p" '$4 ~ pat {print $NF}' | grep -oP 'pid=\K[0-9]+' | sort -u); do + kill -9 "$pid" 2>/dev/null + done + sleep 1 +} + +kill_group() { + local pgid="$1" + kill -TERM -- "-$pgid" 2>/dev/null + for _ in $(seq 10); do + pgrep -g "$pgid" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$pgid" 2>/dev/null + sleep 1 +} + +report_strays() { + local what="$1" pattern="$2" pids + pids=$(pgrep -d, -f "$pattern") + if [ -n "$pids" ]; then + log "$what: $(echo "$pids" | tr ',' '\n' | wc -l) process(es) alive" + ps -o pid,ppid,pgid,etimes,args -p "$pids" | cut -c1-180 + return 1 + fi + return 0 +} + +source_manifest() { + local py="$1" pkg="$2" dir + dir=$("$py" -c "import importlib.util as u; print(u.find_spec('$pkg').submodule_search_locations[0])" 2>/dev/null | tail -1) + if [ -z "$dir" ] || [ ! -d "$dir" ]; then + echo "unavailable $pkg" + return + fi + (cd "$dir" && find . -name '*.py' -not -path '*/__pycache__/*' | sort | while read -r f; do + printf '%s %s/%s\n' "$(tr -d '\r' < "$f" | sha256sum | cut -d' ' -f1)" "$pkg" "${f#./}" + done) +} + +log "cell=$CELL policy=$POLICY_LABEL suite=$SUITE task_id=$TASK_ID episodes=$NEP procs=$NPROC port=$PORT budget=${BUDGET_S}s checkpoint=${CKPT_DIR:-none} res=$RES" + +if [ -n "$CKPT_DIR" ] && [ ! -f "$CKPT_DIR/model.safetensors" ]; then + log "no model.safetensors in $CKPT_DIR" + exit 2 +fi + +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif ! report_strays "refusing to start, env client processes" "$CLIENT_PATTERN" \ + || ! report_strays "refusing to start, pi0 server processes" "$SERVER_PATTERN"; then + exit 2 +fi + +rm -rf "$OUT" +mkdir -p "$OUT" + +"$UV" pip freeze --python "$SPY" > "$RES/environment-server.txt" 2>/dev/null +"$UV" pip freeze --python "$CPY" > "$RES/environment-client.txt" 2>/dev/null +nvidia-smi --query-gpu=index,name,memory.total,driver_version --format=csv > "$RES/environment-gpu.txt" 2>/dev/null +{ + PYTHONPATH="$R/plugrl-server-e32/src" source_manifest "$SPY" plugrl_server + source_manifest "$CPY" plugrl_env_client +} > "$RES/environment-source-$CELL.txt" +log "source manifest: $(wc -l < "$RES/environment-source-$CELL.txt") files, sha256 $(sha256sum "$RES/environment-source-$CELL.txt" | cut -c1-16)" +if [ -n "$CKPT_DIR" ]; then + log "checkpoint sha256: $(sha256sum "$CKPT_DIR/model.safetensors" | cut -c1-16)" +fi + +CKPT_ARGS=() +if [ -n "$CKPT_DIR" ]; then + CKPT_ARGS=(--algo.policy-checkpoint-path "$CKPT_DIR") +fi + +free_port "$PORT" + +( + cd "$R/plugrl-server-e32" && exec setsid env \ + PYTHONPATH="$R/plugrl-server-e32/src" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" CUDA_VISIBLE_DEVICES=${EVAL_SRV_GPU:-0} \ + "$SPY" -m plugrl_server.cli pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda \ + "${CKPT_ARGS[@]}" \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + free_port "$PORT" + if report_strays "after teardown, env client processes" "$CLIENT_PATTERN" \ + && report_strays "after teardown, pi0 server processes" "$SERVER_PATTERN"; then + log "teardown clean" + else + log "TEARDOWN_NOT_CLEAN" + fi +} + +# Cold, after hours out of Lustre's cache, the server has taken over 600 s to +# import and construct. +for _ in $(seq 1500); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened" + tail -20 "$OUT/server.log" + teardown + exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${EVAL_CLI_GPU:-1} MUJOCO_EGL_DEVICE_ID=${EVAL_CLI_GPU:-1} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout "$BUDGET_S" "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs "$NPROC" --num-episodes "$NEP" \ + --env.task-suite-name "$SUITE" \ + --env.task-id "$TASK_ID" \ + --env.no-randomize-initial-state \ + --runner.replan-steps 5 \ + --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +SERVER_DIED=0 +while alive "$CPID"; do + if ! alive "$SPID"; then + log "server exited while the client was running" + SERVER_DIED=1 + kill_group "$CPID" + break + fi + sleep 5 +done +wait "$CPID" +CEXIT=$? +[ "$SERVER_DIED" = 1 ] && [ "$CEXIT" = 0 ] && CEXIT=97 +WALL=$(( $(date +%s) - T0 )) +log "client exited with $CEXIT after ${WALL}s" + +teardown + +"$SPY" "$R/e11/e11_stageC_eval_post.py" --cell "$CELL" --policy-label "$POLICY_LABEL" --task-id "$TASK_ID" \ + --nep "$NEP" --out "$OUT" --res "$RES" --client-exit "$CEXIT" --wall "$WALL" +log "STAGEC_EVAL_DONE $CELL" diff --git a/experiments/e36-pi0-longer/movement.py b/experiments/e36-pi0-longer/movement.py new file mode 100644 index 0000000..6ec115d --- /dev/null +++ b/experiments/e36-pi0-longer/movement.py @@ -0,0 +1,81 @@ +"""How far did each arm's iteration move the policy? Compare checkpoints with the base. + + python movement.py BASE CHECKPOINT [CHECKPOINT ...] + +E26's check_frozen.py with one line added. Per module group - `mlp`, `attn`, +`mod` and `io`, E22's partition - how many tensors moved and the relative +distance; then `expert`, the relative distance over every tensor of the +action expert (the 201 whose name contains `gemma_expert`), which is E15's +`actor.action_expert` row: 0.0151 for E14's iteration, 0.0066 for the +smallest movement E15 found destructive. +""" + +import collections +import re +import sys + +import torch +from safetensors import safe_open + +GROUPS = [ + ( + "mod", + re.compile( + r"(input_layernorm|post_attention_layernorm|model\.norm)\.dense\.(weight|bias)$" + ), + ), + ("attn", re.compile(r"self_attn\.(q|k|v|o)_proj\.weight$")), + ("mlp", re.compile(r"mlp\.(gate|up|down)_proj\.weight$")), + ( + "io", + re.compile( + r"^actor\.(action_in_proj|action_out_proj|time_mlp_in|time_mlp_out)\.(weight|bias)$" + ), + ), +] + + +def group_of(key: str) -> str | None: + if key.startswith("actor.paligemma_with_expert.") and "gemma_expert" not in key: + return None + hits = [name for name, rx in GROUPS if rx.search(key)] + return hits[0] if len(hits) == 1 else None + + +def is_expert(key: str) -> bool: + return "gemma_expert" in key + + +def main(base_path: str, ckpts: list[str]) -> None: + with safe_open(base_path, framework="pt", device="cpu") as fb: + keys = [k for k in fb.keys() if group_of(k) or is_expert(k)] + base = {k: fb.get_tensor(k) for k in keys} + for path in ckpts: + moved = collections.Counter() + still = collections.Counter() + num = collections.defaultdict(float) + den = collections.defaultdict(float) + with safe_open(path, framework="pt", device="cpu") as fc: + for k in keys: + a, b = base[k].double(), fc.get_tensor(k).double() + d2, n2 = float((b - a).pow(2).sum()), float(a.pow(2).sum()) + labels = [g for g in (group_of(k),) if g] + if is_expert(k): + labels.append("expert") + for g in labels: + if torch.equal(a, b): + still[g] += 1 + else: + moved[g] += 1 + num[g] += d2 + den[g] += n2 + print(path) + for g in [name for name, _ in GROUPS] + ["expert"]: + print( + f" {g:6s} moved {moved[g]:3d} unchanged {still[g]:3d} " + f"relative distance {(num[g] / den[g]) ** 0.5:.6f}" + ) + + +if __name__ == "__main__": + main(sys.argv[1], sys.argv[2:]) diff --git a/experiments/e36-pi0-longer/run.sh b/experiments/e36-pi0-longer/run.sh new file mode 100644 index 0000000..980a11b --- /dev/null +++ b/experiments/e36-pi0-longer/run.sh @@ -0,0 +1,117 @@ +#!/usr/bin/env bash +# E36 on qz103. See PROTOCOL.md. +# +# setsid nohup bash e36/run.sh > e36/run.out 2>&1 < /dev/null & +# +# Phase 1, in parallel: E32's fpopp for ten iterations on two server seeds, +# and E25's DPPO for ten. Phase 2, in parallel: every run's iteration-5 and +# iteration-10 checkpoints evaluated; then movements and learn statistics. +# +# Cards, with the EGL mapping measured for E32 (EGL id -> nvidia-smi index +# 0->3 1->2 2->0 3->1 4->7 5->6 6->4 7->5): each run's clients render on its +# own second card, and each evaluation's single client on its server's card. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +E=$R/e36 +LOG=$E/run.log +mkdir -p "$E/results" + +log() { echo "[$(date '+%m-%d %H:%M:%S')] $*" | tee -a "$LOG"; } + +nth_ckpt() { # ROOT N: the N-th checkpoint directory under ROOT, by step + local step + step=$(ls -1 "$1" 2>/dev/null | grep -E '^[0-9]+$' | sort -n | sed -n "${2}p") + [ -n "$step" ] && echo "$1/$step" +} +root_of() { # RUN: where its checkpoints are + case "$1" in + dppo) echo "$E/runs/dppo/ck/dppo/pi0-policy/dppo" ;; + *) echo "$E/$1/ck/fpo/pi0-policy/$1" ;; + esac +} + +evaluate() { # RUN ITERATION SRV_GPU CLI_GPU PORT + local ck + ck=$(nth_ckpt "$(root_of "$1")" "$2") + if [ -z "$ck" ]; then + log "END eval $1-it$2 not run: no checkpoint" + return + fi + ALLOW_SIBLINGS=1 EVAL_SRV_GPU="$3" EVAL_CLI_GPU="$4" \ + bash "$E/eval.sh" "e36-$1-it$2" "$1-it$2" 8 50 "$5" 21600 "$ck" > "$E/eval-$1-it$2.out" 2>&1 + log "END eval $1-it$2 rc=$?" +} + +FPOPP="--algo.n-critic-warmup-itrs 1 --algo.output-mode u --algo.no-discretize-t-for-training --algo.cfm-loss-steps 5 --algo.cfm-loss-dims 7 --algo.cfm-loss-sum-over-steps --algo.ratio-per-sample" + +log "code: $(cat $R/plugrl-server-e32/COMMIT)" +log "--- phase 1" +( + ALLOW_SIBLINGS=1 SEED=7 SRV_GPUS=0,1 CLI_GPU=3 ARM_FLAGS="$FPOPP" \ + bash "$E/train.sh" fpopp-s7 10 8471 cuda:1 > "$E/train-fpopp-s7.out" 2>&1 + log "END train fpopp-s7 rc=$?" +) & +A=$! +sleep 60 +( + ALLOW_SIBLINGS=1 SEED=8 SRV_GPUS=2,3 CLI_GPU=0 ARM_FLAGS="$FPOPP" \ + bash "$E/train.sh" fpopp-s8 10 8472 cuda:1 > "$E/train-fpopp-s8.out" 2>&1 + log "END train fpopp-s8 rc=$?" +) & +B=$! +sleep 60 +( + ALLOW_SIBLINGS=1 SRV_GPUS=4,5 CLI_GPU=7 PORT=8473 \ + bash "$E/dppo_cell.sh" dppo 10 4096 > "$E/train-dppo.out" 2>&1 + log "END train dppo rc=$?" +) & +C=$! +wait "$A" "$B" "$C" + +: > "$E/results/configs.txt" +for run in fpopp-s7 fpopp-s8; do + { echo "## $run"; grep -h "Algorithm: fpo" "$E/$run/server.log" | tail -1; } >> "$E/results/configs.txt" +done +{ echo "## dppo"; grep -h "Algorithm: dppo" "$E/runs/dppo/server.log" | tail -1; } >> "$E/results/configs.txt" +for run in fpopp-s7 fpopp-s8 dppo; do + log "checkpoints $run: $(ls -1 "$(root_of $run)" | grep -E '^[0-9]+$' | sort -n | tr '\n' ' ')" +done + +log "--- phase 2" +evaluate fpopp-s7 5 0 2 8571 & +P1=$! +evaluate fpopp-s7 10 1 3 8572 & +P2=$! +evaluate fpopp-s8 5 2 1 8573 & +P3=$! +evaluate fpopp-s8 10 3 0 8574 & +P4=$! +evaluate dppo 5 4 6 8575 & +P5=$! +evaluate dppo 10 5 7 8576 & +P6=$! + +log "--- movement from the base, iterations 5 and 10" +CKS=() +for run in fpopp-s7 fpopp-s8 dppo; do + for it in 5 10; do + ck=$(nth_ckpt "$(root_of $run)" $it) && CKS+=("$ck/model.safetensors") + done +done +"$R/venv/bin/python" "$E/movement.py" "$R/e14/base-statedict/model.safetensors" \ + "${CKS[@]}" > "$E/results/movement.txt" 2>&1 +tee -a "$LOG" < "$E/results/movement.txt" + +log "--- learn-step statistics" +for run in fpopp-s7 fpopp-s8 dppo; do + echo "## $run" + if [ "$run" = dppo ]; then d="$E/runs/dppo"; else d="$E/$run"; fi + "$R/venv/bin/python" "$E/tb_read.py" "$d" 2>&1 | grep -v Warning +done > "$E/results/learn-stats.txt" + +wait "$P1" "$P2" "$P3" "$P4" "$P5" "$P6" + +log "=== rows ===" +tee -a "$LOG" < "$E/results/stageC_eval.tsv" +log "E36_DONE" diff --git a/experiments/e36-pi0-longer/tb_read.py b/experiments/e36-pi0-longer/tb_read.py new file mode 100644 index 0000000..cf6389d --- /dev/null +++ b/experiments/e36-pi0-longer/tb_read.py @@ -0,0 +1,33 @@ +"""Print a run's training curves from its tensorboards: every iteration's value. + + python tb_read.py ROOT [ROOT ...] + +For each events file under a root: the rollout success, return and length, +and the algorithm's ratio, clip, KL and CFM-loss scalars, one line per tag +with every iteration's value. +""" + +import glob +import sys + +from tensorboard.backend.event_processing.event_accumulator import EventAccumulator + +KEYS = ( + "rollout/success", + "rollout/reward", + "rollout/length", + "ratio_mean", + "clipped_ratio", + "cfm_loss_mean", + "approx_kl", + "clipfrac", +) +for root in sys.argv[1:]: + for ev in sorted(glob.glob(root + "/**/events.*", recursive=True)): + acc = EventAccumulator(ev, size_guidance={"scalars": 0}) + acc.Reload() + print(ev) + for t in acc.Tags()["scalars"]: + if any(k in t for k in KEYS): + values = " ".join(f"{e.value:.4g}" for e in acc.Scalars(t)) + print(f" {t:36s} {values}") diff --git a/experiments/e36-pi0-longer/train.sh b/experiments/e36-pi0-longer/train.sh new file mode 100644 index 0000000..6e2a35a --- /dev/null +++ b/experiments/e36-pi0-longer/train.sh @@ -0,0 +1,214 @@ +#!/usr/bin/env bash +# Does a second learn step fit once the optimizer's state is on another card, +# and what does the transfer cost? +# +# bash e36/train.sh CELL ITERATIONS PORT [MASTER_DEVICE] +# +# E36: E32's training harness with its output under e36/ and the +# server's seed taken from SEED (default 7, E14's). E32's header follows. +# +# bash e32/train.sh CELL ITERATIONS PORT [MASTER_DEVICE] +# +# E32: E14's training harness, unchanged but for three things - the +# server runs $R/plugrl-server-e32 (main plus #74, deployed LF) through +# PYTHONPATH, output goes under e32/, and ARM_FLAGS, when set, is split +# on spaces and passed to the server after every other flag. +# +# E11 stopped at one iteration of ten: the float32 copies and Adam's moments +# do not exist until the first learn step and then have to fit beside +# everything the second one needs. Measured, that state is 3,578 MiB; moving +# it to a second card leaves the model's card with 8,046 MiB instead of +# 11,624. +# +# Two iterations is enough to answer the question, since it was the second +# that died. Ten is the experiment, and only worth starting if this passes. +set -uo pipefail + +R=/home/gotham/tmp/plugrl +CELL="$1" +ITERATIONS="${2:-2}" +PORT="${3:-8161}" +MASTER_DEVICE="${4:-cuda:1}" + +BUFFER_SIZE=${BUFFER_SIZE:-4096} +BATCH_SIZE=8 +N_SAMPLES=4 +TASK_ID=8 +SUITE=libero_10 +NPROC=10 +GLOBAL_STEPS=$(( ITERATIONS * BUFFER_SIZE )) + +SPY=$R/venv/bin/python +CPY=$R/venv-libero/bin/python +OUT=$R/e36/$CELL + +CLIENT_PATTERN="[v]env-libero.*/bin/python" +SERVER_PATTERN="[p]lugrl_server.cli pi0-policy" + +# The memory recorder samples by physical index, which CUDA_VISIBLE_DEVICES +# does not remap. Follow the cards the run was actually given. +_SG="${SRV_GPUS:-0,1}" +export MEM_G0="${_SG%%,*}" +export MEM_G1="${_SG##*,}" + +log() { echo "[$(date +%H:%M:%S)] $*"; } + +alive() { + local s + s=$(ps -o stat= -p "$1" 2>/dev/null | tr -d ' ') + [ -n "$s" ] && [ "${s#Z}" = "$s" ] +} + +kill_group() { + local pgid="$1" + kill -TERM -- "-$pgid" 2>/dev/null + for _ in $(seq 30); do + pgrep -g "$pgid" > /dev/null 2>&1 || return 0 + sleep 1 + done + kill -KILL -- "-$pgid" 2>/dev/null + sleep 1 +} + +report_strays() { + local what="$1" pattern="$2" pids + pids=$(pgrep -f "$pattern" | while read -r p; do + [ "$(ps -o comm= -p "$p" 2>/dev/null)" = python ] && echo "$p" + done | paste -sd, -) + if [ -n "$pids" ]; then + log "$what: $(echo "$pids" | tr ',' '\n' | wc -l) alive" + ps -o pid,etimes,args -p "$pids" | cut -c1-120 | head -5 + return 1 + fi + return 0 +} + +log "cell=$CELL iterations=$ITERATIONS master_device=$MASTER_DEVICE lr=${LR:-1e-5} updates=${UPDATES:-4} clip=${CLIP_EPS:-0.05} drift=${DRIFT:-0} warmup=${WARMUP:-0} srv_gpus=${SRV_GPUS:-0,1} cli_gpu=${CLI_GPU:-2} buffer=$BUFFER_SIZE batch=$BATCH_SIZE" + +if [ "${ALLOW_SIBLINGS:-0}" = 1 ]; then + log "ALLOW_SIBLINGS=1: not checking for other clients or servers" +elif ! report_strays "refusing to start, clients" "$CLIENT_PATTERN" \ + || ! report_strays "refusing to start, servers" "$SERVER_PATTERN"; then + exit 2 +fi + +rm -rf "$OUT"; mkdir -p "$OUT" + +MASTER_FLAG=() +[ "$MASTER_DEVICE" != "none" ] && MASTER_FLAG=(--algo.master-weights-device "$MASTER_DEVICE") + +# Only passed when set, so the harness still runs against a server that does +# not have the flag at all. +OPT_FLAGS=() +[ -n "${WARMUP:-}" ] && [ "${WARMUP:-0}" != 0 ] \ + && OPT_FLAGS+=(--algo.critic-warmup-iterations "$WARMUP") +[ -n "${DRIFT:-}" ] && [ "${DRIFT:-0}" != 0 ] \ + && OPT_FLAGS+=(--algo.max-policy-drift "$DRIFT") +if [ -n "${ARM_FLAGS:-}" ]; then + read -r -a _ARM <<< "$ARM_FLAGS" + OPT_FLAGS+=("${_ARM[@]}") +fi +log "arm flags: ${ARM_FLAGS:-none}" + +( + cd "$R/plugrl-server-e32" && exec setsid env \ + PYTHONPATH="$R/plugrl-server-e32/src" \ + OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" \ + CUDA_VISIBLE_DEVICES=${SRV_GPUS:-0,1} \ + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \ + "$SPY" -m plugrl_server.cli pi0-policy default fpo default \ + --policy.name pi05_libero \ + --policy.checkpoint-path "$R/ckpt/pi05_libero" \ + --policy.device cuda:0 \ + --algo.learning-rate ${LR:-1e-5} \ + --algo.batch-size "$BATCH_SIZE" \ + --algo.num-updates-per-batch ${UPDATES:-4} \ + --algo.n-samples-per-action "$N_SAMPLES" \ + --algo.buffer-size "$BUFFER_SIZE" \ + --algo.clipping-epsilon ${CLIP_EPS:-0.05} \ + --algo.global-steps "$GLOBAL_STEPS" \ + --algo.save-interval 1 \ + "${MASTER_FLAG[@]}" \ + "${OPT_FLAGS[@]}" \ + --seed ${SEED:-7} \ + --port "$PORT" \ + --no-show-progress-bar --no-show-metric-table \ + --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite +) > "$OUT/server.log" 2>&1 & +SPID=$! + +# Peak memory on both cards, sampled from outside: nothing in the server +# records it, and which card holds what is the point here. +( + exec setsid bash -c 'while true; do echo "$(date +%s),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G0),$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits -i $MEM_G1)"; sleep 10; done' +) > "$OUT/gpu_mem.csv" 2>/dev/null & +MPID=$! + +teardown() { + kill_group "$SPID" + [ -n "${CPID:-}" ] && kill_group "$CPID" + kill_group "$MPID" + if report_strays "after teardown, clients" "$CLIENT_PATTERN" \ + && report_strays "after teardown, servers" "$SERVER_PATTERN"; then + log "teardown clean" + else + log "TEARDOWN_NOT_CLEAN" + fi +} + +for _ in $(seq 1800); do + grep -q "is listening" "$OUT/server.log" 2>/dev/null && break + alive "$SPID" || break + sleep 1 +done +if ! grep -q "is listening" "$OUT/server.log" 2>/dev/null; then + log "server never listened"; tail -15 "$OUT/server.log"; teardown; exit 1 +fi +log "server listening" + +T0=$(date +%s) +( + cd "$OUT" && exec setsid env \ + LIBERO_CONFIG_PATH="$R/.libero" MUJOCO_GL=egl PYOPENGL_PLATFORM=egl \ + CUDA_VISIBLE_DEVICES=${CLI_GPU:-2} MUJOCO_EGL_DEVICE_ID=${CLI_GPU:-2} \ + XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" \ + timeout 43200 "$CPY" -m plugrl_env_client.cli libero-v1 \ + --server-host 127.0.0.1 --server-port "$PORT" \ + --num-envs 1 --num-procs "$NPROC" --num-episodes 1000000000 \ + --env.task-suite-name "$SUITE" \ + --env.task-id "$TASK_ID" \ + --env.randomize-initial-state \ + --runner.replan-steps 5 --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name "$CELL" +) > "$OUT/client.log" 2>&1 & +CPID=$! + +SERVER_DIED=0 +while alive "$CPID"; do + if ! alive "$SPID"; then + log "server exited" + SERVER_DIED=1 + for _ in $(seq 60); do alive "$CPID" || break; sleep 1; done + kill_group "$CPID" + break + fi + sleep 15 +done +wait "$CPID" +CEXIT=$? +WALL=$(( $(date +%s) - T0 )) +log "clients exit $CEXIT after ${WALL}s (server_died=$SERVER_DIED)" + +teardown + +log "=== learn steps seen ===" +grep -cE "learn|Learn" "$OUT/server.log" 2>/dev/null | head -1 +log "=== out of memory? ===" +grep -c "OutOfMemoryError" "$OUT/server.log" 2>/dev/null | head -1 +log "=== peak memory, GPU0 and GPU1 ===" +awk -F, 'NR>0 {if ($2>m0) m0=$2; if ($3>m1) m1=$3} END {print "gpu0_peak_mib=" m0 " gpu1_peak_mib=" m1}' "$OUT/gpu_mem.csv" 2>/dev/null +log "=== checkpoints ===" +find "$OUT/ck" -name model.safetensors -printf '%TH:%TM %s %p\n' 2>/dev/null | sort | tail -5 +log "E36_TRAIN_DONE $CELL" From 4887dcc6eadf01451a7cd2e679761019f1a966d0 Mon Sep 17 00:00:00 2001 From: tactino <18781106300@163.com> Date: Sun, 27 Sep 2026 22:25:34 -0400 Subject: [PATCH 2/2] exp: E36 results - FPO++ collapses pi0.5 within five iterations, DPPO holds it at 20/50 by barely moving it --- experiments/e36-pi0-longer/FINDINGS.md | 93 ++++++++++ .../e36-pi0-longer/results/configs.txt | 6 + .../results/environment-client.txt | 132 ++++++++++++++ .../results/environment-gpu.txt | 9 + .../results/environment-server.txt | 165 ++++++++++++++++++ .../environment-source-e36-dppo-it10.txt | 109 ++++++++++++ .../environment-source-e36-dppo-it5.txt | 109 ++++++++++++ .../environment-source-e36-fpopp-s7-it5.txt | 109 ++++++++++++ .../environment-source-e36-fpopp-s8-it5.txt | 109 ++++++++++++ .../e36-pi0-longer/results/learn-stats.txt | 26 +++ .../e36-pi0-longer/results/movement.txt | 24 +++ experiments/e36-pi0-longer/results/run.log | 48 +++++ experiments/e36-pi0-longer/results/run.out | 48 +++++ .../e36-pi0-longer/results/stageC_eval.tsv | 5 + .../results/stageC_eval_correctness.tsv | 5 + .../e36-pi0-longer/results/train-dppo.out | 8 + .../e36-pi0-longer/results/train-fpopp-s7.out | 26 +++ .../e36-pi0-longer/results/train-fpopp-s8.out | 19 ++ 18 files changed, 1050 insertions(+) create mode 100644 experiments/e36-pi0-longer/FINDINGS.md create mode 100644 experiments/e36-pi0-longer/results/configs.txt create mode 100644 experiments/e36-pi0-longer/results/environment-client.txt create mode 100644 experiments/e36-pi0-longer/results/environment-gpu.txt create mode 100644 experiments/e36-pi0-longer/results/environment-server.txt create mode 100644 experiments/e36-pi0-longer/results/environment-source-e36-dppo-it10.txt create mode 100644 experiments/e36-pi0-longer/results/environment-source-e36-dppo-it5.txt create mode 100644 experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s7-it5.txt create mode 100644 experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s8-it5.txt create mode 100644 experiments/e36-pi0-longer/results/learn-stats.txt create mode 100644 experiments/e36-pi0-longer/results/movement.txt create mode 100644 experiments/e36-pi0-longer/results/run.log create mode 100644 experiments/e36-pi0-longer/results/run.out create mode 100644 experiments/e36-pi0-longer/results/stageC_eval.tsv create mode 100644 experiments/e36-pi0-longer/results/stageC_eval_correctness.tsv create mode 100644 experiments/e36-pi0-longer/results/train-dppo.out create mode 100644 experiments/e36-pi0-longer/results/train-fpopp-s7.out create mode 100644 experiments/e36-pi0-longer/results/train-fpopp-s8.out diff --git a/experiments/e36-pi0-longer/FINDINGS.md b/experiments/e36-pi0-longer/FINDINGS.md new file mode 100644 index 0000000..e969f63 --- /dev/null +++ b/experiments/e36-pi0-longer/FINDINGS.md @@ -0,0 +1,93 @@ +# E36: over more iterations FPO++ runs pi0.5 into the ground, and DPPO barely moves it + +2026-09-27/28 · qz103, two cards per run · pi0.5 on LIBERO-10 task 8 · +protocol: [`PROTOCOL.md`](PROTOCOL.md) (`bb9d36b`, before any run) + +--- + +## The result + +Fifty-episode evaluations, initial states 0-49 (the untrained policy: 28-37 +over seven evaluations, mean 30.9, sd 3.6): + +| run | iteration 5 | iteration 10 | +| --- | --- | --- | +| `fpopp-s7` (FPO with FPO++'s chunk loss and per-sample ratio) | **5** / 50 | not run: no checkpoint | +| `fpopp-s8` | **0** / 50 | not run: no checkpoint | +| `dppo` (the `libero` variant) | 28 / 50 | **20** / 50 | + +* **V1 and V2 pass.** Every run's flags took (`results/configs.txt`). Every + evaluation ran fifty episodes, its client exited 0, and it is valid + (`results/stageC_eval_correctness.tsv`). +* **P1 fails for both FPO++ runs**, and the fault is mine. The training + harness, derived from E32's, wraps the clients in E32's `timeout 43200`. + That fits two iterations. Ten FPO iterations of about 80 minutes do not + fit. At 12 hours the clients were killed and the server with them, in the + middle of the tenth learn step (`results/run.out`; the server log's + "Received SIGTERM ... the model lock was still held"). So nine iterations + learned and no iteration-10 checkpoint exists. DPPO's ten iterations of + about an hour fit, and it ran to completion. +* **P2 (FPO++ holds at iteration 10) cannot be read as registered.** Its + checkpoint does not exist. The evidence it would have read points one way: + - At iteration 5 both runs are at the collapse line or below it (5 and 0). + - The training rollouts had fallen to zero: 0.05, 0, 0 and 0, 0, 0 over + iterations 7-9. + - The one restart the protocol allows was not used. It would cost about + thirteen GPU-hours to read a number the rest already gives. +* **P3 (DPPO holds at iteration 10) holds, on the line**: 20 of 50, where + holding is 20 or more. That is also three standard deviations below the + untrained policy's mean. It went 28 at iteration 5, then 20. + +Neither algorithm made pi0.5 better. The status rule's "learns" is 42 of 50 +at iteration 10, and nothing came near it. `pi0-policy` · DPPO stays +"holds", now after ten iterations. `pi0-policy` · FPO held for one update +(E32) and collapsed within five. + +--- + +## What the logs show + +`results/learn-stats.txt`, per iteration: + +| | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | +| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | +| s7 training success | 0.60 | 0.38 | 0.73 | 0.40 | 0.10 | 0.14 | 0.05 | 0 | 0 | +| s7 CFM loss of its own samples | 0.11 | 0.13 | 0.16 | 0.26 | 0.81 | 5.8 | 28 | 74 | 114 | +| s7 clipped fraction | 0 | 0.46 | 0.52 | 0.46 | 0.43 | 0.50 | 0.58 | 0.75 | 0.96 | +| s8 training success | 0.67 | 0.50 | 0.47 | 0.18 | 0.13 | 0 | 0 | 0 | 0 | +| s8 CFM loss of its own samples | 0.12 | 0.12 | 0.19 | 1.9 | 10 | 52 | 180 | 338 | 407 | + +The flow-matching loss of the policy's own freshly sampled actions is the +loss the next ratio starts from. It rose a thousandfold, and the success +fell with it. The velocity field stopped describing the distribution it +samples: the policy came apart, it did not merely drift. The clipped +fraction climbing towards one says that by the end almost every sample was +outside the trust region the moment the update began. + +DPPO's `approx_kl` stayed between 1.3 and 4.1 x 10⁻⁷ for all ten +iterations, and the action expert moved 0.16% by iteration 5 and 0.24% by +iteration 10 (FPO++: 2.8% by iteration 5; `results/movement.txt`). DPPO kept +pi0.5 because it hardly touched it, as in E25. Where it did end up, 20 of +50, is below where it started. + +--- + +## Where this points + +The same shape as E37 on square: FPO's first update is harmless, later ones +degrade a pretrained policy. E37's critic never fit its returns. Here the +flow model's loss on its own samples runs away. Both ran with the parts of +FPO's own defaults that FPO++ does not share (`results/configs.txt`): +- rewards scaled by 10 +- values and advantages recomputed every epoch +- one learning rate for actor and critic +- no gradient clipping + +E39 is running FPO++'s +square fine-tuning in full (#91) to see whether the rest of FPO++'s setting +stops the square fall. If it does, the same settings go to pi0.5 next. For +DPPO on pi0.5 the question is the other one, why it barely moves, and +fpo-policy's history (E29, E30: the noise level) is the first place to look. + +Checkpoints and tensorboards stay on qz103. The results, logs and harness +outputs are here. diff --git a/experiments/e36-pi0-longer/results/configs.txt b/experiments/e36-pi0-longer/results/configs.txt new file mode 100644 index 0000000..1eeea26 --- /dev/null +++ b/experiments/e36-pi0-longer/results/configs.txt @@ -0,0 +1,6 @@ +## fpopp-s7 +13:28:56|INFO|Algorithm: fpo, Config: FPOAlgoConfig(global_steps=40960, buffer_size=4096, output_mode='u', fpo_playground_trick=True, treat_truncated_as_done=False, discounting=0.995, reward_scaling=10.0, gae_lambda=0.95, batch_size=8, num_updates_per_batch=4, learning_rate=1e-05, value_loss_coeff=0.25, clipping_epsilon=0.05, normalize_advantage=True, n_critic_warmup_itrs=1, max_policy_drift=0.0, n_samples_per_action=4, discretize_t_for_training=False, average_losses_before_exp=True, cfm_loss_steps=5, cfm_loss_dims=7, cfm_loss_sum_over_steps=True, ratio_per_sample=True, save_interval=1, policy_checkpoint_path=None, restore='all', master_weights_device='cuda:1') +## fpopp-s8 +13:29:52|INFO|Algorithm: fpo, Config: FPOAlgoConfig(global_steps=40960, buffer_size=4096, output_mode='u', fpo_playground_trick=True, treat_truncated_as_done=False, discounting=0.995, reward_scaling=10.0, gae_lambda=0.95, batch_size=8, num_updates_per_batch=4, learning_rate=1e-05, value_loss_coeff=0.25, clipping_epsilon=0.05, normalize_advantage=True, n_critic_warmup_itrs=1, max_policy_drift=0.0, n_samples_per_action=4, discretize_t_for_training=False, average_losses_before_exp=True, cfm_loss_steps=5, cfm_loss_dims=7, cfm_loss_sum_over_steps=True, ratio_per_sample=True, save_interval=1, policy_checkpoint_path=None, restore='all', master_weights_device='cuda:1') +## dppo +13:30:55|INFO|Algorithm: dppo, Config: DPPOAlgoConfigLibero(global_steps=40960, gamma=0.999, gamma_denoising=1.0, actor_lr=5e-06, critic_lr=0.0001, actor_weight_decay=0.0, critic_weight_decay=0.0, actor_lr_scheduler=None, critic_lr_scheduler=None, buffer_size=4096, gae_lambda=0.95, update_epochs=4, norm_adv=True, clip_ploss_coef=0.001, clip_ploss_coef_base=0.001, clip_ploss_coef_rate=3, clip_vloss_coef=inf, ent_coef=0.0, vf_coef=0.5, max_grad_norm=inf, target_kl=inf, logprob_noise_level=0.5, sampling_noise_level=0.5, logprob_clamp_min=-5.0, logprob_clamp_max=2.0, clip_advantage_lower_quantile=0, clip_advantage_upper_quantile=1, n_critic_warmup_itrs=0, use_normalized_rewards=False, batch_size=8, train_itrs=10, save_interval=1, grad_accum_steps=16, policy_checkpoint_path=None, restore='all') diff --git a/experiments/e36-pi0-longer/results/environment-client.txt b/experiments/e36-pi0-longer/results/environment-client.txt new file mode 100644 index 0000000..1c71c81 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-client.txt @@ -0,0 +1,132 @@ +absl-py==2.5.0 +annotated-doc==0.0.5 +antlr4-python3-runtime==4.9.3 +anyio==4.15.1 +attrs==26.1.0 +bddl==3.6.0 +certifi==2026.7.22 +click==8.5.0 +cloudpickle==3.1.2 +contourpy==1.3.3 +cuda-bindings==13.4.1 +cuda-pathfinder==1.8.1 +cuda-toolkit==13.0.3.0 +cycler==0.12.1 +defusedxml==0.7.1 +docstring-parser==0.18.0 +easydict==1.13 +egl-probe==1.0.2 +einops==0.8.2 +etils==1.14.0 +evdev==2.0.0 +farama-notifications==0.0.6 +fastjsonschema==2.22.2 +filelock==3.32.6 +fonttools==4.65.0 +fsspec==2026.7.0 +future==1.0.0 +glfw==2.10.2 +grpcio==1.83.1 +gym==0.26.2 +gym-notices==0.1.0 +gymnasium==1.3.0 +h11==0.16.0 +h5py==3.16.0 +hf-xet==1.6.0 +httpcore==1.0.9 +httpx==0.28.1 +huggingface-hub==1.31.0 +hydra-core==1.3.6 +idna==3.19 +imageio==2.37.4 +imageio-ffmpeg==0.6.0 +iniconfig==2.3.0 +jinja2==3.1.6 +joblib==1.6.0 +jsonschema==4.26.0 +jsonschema-specifications==2025.9.1 +jupyter-core==5.9.1 +jupytext==1.19.5 +kiwisolver==1.5.1 +-e file:///home/gotham/tmp/plugrl/LIBERO +llvmlite==0.49.0 +loguru==0.7.3 +markdown==3.10.3 +markdown-it-py==4.2.0 +markupsafe==3.0.3 +matplotlib==3.11.2 +mdit-py-plugins==0.6.1 +mdurl==0.1.2 +mpmath==1.3.0 +msgpack==1.2.2 +mujoco==2.3.7 +nbformat==5.11.1 +networkx==3.6.1 +nltk==3.10.3 +numba==0.67.0 +numpy==1.26.4 +nvidia-cublas==13.1.1.3 +nvidia-cuda-cupti==13.0.85 +nvidia-cuda-nvrtc==13.0.88 +nvidia-cuda-runtime==13.0.96 +nvidia-cudnn-cu13==9.24.0.43 +nvidia-cufft==12.0.0.61 +nvidia-cufile==1.15.1.6 +nvidia-curand==10.4.0.35 +nvidia-cusolver==12.0.4.66 +nvidia-cusparse==12.6.3.3 +nvidia-cusparselt-cu13==0.8.1 +nvidia-nccl-cu13==2.30.7 +nvidia-nvjitlink==13.4.52 +nvidia-nvshmem-cu13==3.4.5 +nvidia-nvtx==13.0.85 +omegaconf==2.3.1 +opencv-python==4.11.0.86 +opencv-python-headless==4.11.0.86 +packaging==26.3 +pillow==12.3.0 +platformdirs==4.11.8 +pluggy==1.6.0 +-e file:///home/gotham/tmp/plugrl/plugrl-env-client +-e file:///home/gotham/tmp/plugrl/plugrl-protocol +protobuf==7.36.1 +psutil==7.2.2 +pygments==2.21.0 +pynput==1.8.2 +pyopengl==3.1.10 +pyparsing==3.3.2 +pytest==9.1.1 +python-dateutil==2.9.0.post0 +python-xlib==0.33 +pyyaml==6.0.3 +referencing==0.37.0 +regex==2026.9.10 +rich==15.0.0 +robomimic==0.3.0 +robosuite==1.4.1 +rpds-py==2026.6.3 +safetensors==0.8.0 +scipy==1.17.1 +setuptools==84.0.0 +shellingham==1.5.4 +six==1.17.0 +sympy==1.14.0 +tensorboard==2.21.0 +tensorboard-data-server==0.7.2 +tensorboardx==2.6.5 +termcolor==3.3.0 +thop==0.1.1.post2209072238 +tokenizers==0.23.2 +torch==2.14.0 +torchvision==0.29.0 +tqdm==4.70.1 +traitlets==5.16.1 +transformers==5.17.0 +triton==3.8.0 +typeguard==4.6.0 +typer==0.27.2 +typing-extensions==4.16.0 +tyro==1.0.16 +websockets==17.1 +werkzeug==3.1.8 +zipp==4.1.0 diff --git a/experiments/e36-pi0-longer/results/environment-gpu.txt b/experiments/e36-pi0-longer/results/environment-gpu.txt new file mode 100644 index 0000000..1ce8f3d --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-gpu.txt @@ -0,0 +1,9 @@ +index, name, memory.total [MiB], driver_version +0, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +1, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +2, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +3, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +4, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +5, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +6, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 +7, NVIDIA GeForce RTX 3090, 24576 MiB, 580.76.05 diff --git a/experiments/e36-pi0-longer/results/environment-server.txt b/experiments/e36-pi0-longer/results/environment-server.txt new file mode 100644 index 0000000..4480700 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-server.txt @@ -0,0 +1,165 @@ +absl-py==2.5.0 +aiohappyeyeballs==2.7.1 +aiohttp==3.14.3 +aiosignal==1.4.0 +annotated-types==0.8.0 +antlr4-python3-runtime==4.9.3 +attrs==26.1.0 +augmax==0.4.1 +av==18.1.0 +beartype==0.19.0 +beautifulsoup4==4.15.0 +certifi==2026.7.22 +charset-normalizer==3.5.1 +chex==0.1.89 +click==8.5.0 +cloudpickle==3.1.2 +colorama==0.4.6 +contourpy==1.3.3 +cycler==0.12.1 +datasets==3.6.0 +dill==0.3.8 +dm-tree==0.1.10 +docstring-parser==0.18.0 +dppo @ file:///home/gotham/tmp/plugrl/dppo-src +draccus==0.11.6 +einops==0.8.2 +equinox==0.13.8 +etils==1.14.0 +filelock==3.32.6 +flax==0.10.2 +fonttools==4.66.0 +frozenlist==1.8.0 +fsspec==2025.3.0 +gdown==6.4.0 +googleapis-common-protos==1.75.4 +grpcio==1.83.1 +hf-xet==1.6.0 +huggingface-hub==0.36.2 +humanize==4.16.0 +hydra-core==1.3.7 +idna==3.19 +imageio==2.37.4 +importlib-metadata==9.0.1 +iniconfig==2.3.0 +jax==0.5.3 +jaxlib==0.5.3 +jaxtyping==0.2.36 +jinja2==3.1.6 +jsonlines==4.0.0 +jsonschema==4.26.0 +jsonschema-specifications==2025.9.1 +kiwisolver==1.5.1 +-e file:///home/gotham/tmp/plugrl/lerobot +loguru==0.7.3 +markdown==3.10.3 +markdown-it-py==4.2.0 +markupsafe==3.0.3 +matplotlib==3.11.2 +mdurl==0.1.2 +mergedeep==1.3.4 +ml-collections==1.0.0 +ml-dtypes==0.4.1 +mpmath==1.3.0 +msgpack==1.2.2 +multidict==6.8.0 +multiprocess==0.70.16 +mypy-extensions==1.1.0 +nest-asyncio==1.6.0 +networkx==3.6.1 +numpy==1.26.4 +numpydantic==1.8.1 +nvidia-cublas-cu12==12.6.4.1 +nvidia-cuda-cupti-cu12==12.6.80 +nvidia-cuda-nvrtc-cu12==12.6.77 +nvidia-cuda-runtime-cu12==12.6.77 +nvidia-cudnn-cu12==9.5.1.17 +nvidia-cufft-cu12==11.3.0.4 +nvidia-cufile-cu12==1.11.1.6 +nvidia-curand-cu12==10.3.7.77 +nvidia-cusolver-cu12==11.7.1.2 +nvidia-cusparse-cu12==12.5.4.2 +nvidia-cusparselt-cu12==0.6.3 +nvidia-nccl-cu12==2.26.2 +nvidia-nvjitlink-cu12==12.6.85 +nvidia-nvtx-cu12==12.6.77 +omegaconf==2.3.1 +opencv-python-headless==4.11.0.86 +-e file:///home/gotham/tmp/plugrl/plugrl-server/third_party/openpi +-e file:///home/gotham/tmp/plugrl/plugrl-server/third_party/openpi/packages/openpi-client +opentelemetry-api==1.45.0 +opentelemetry-exporter-http-transport==0.66b0 +opentelemetry-exporter-otlp-common==0.66b0 +opentelemetry-exporter-otlp-proto-common==1.45.0 +opentelemetry-exporter-otlp-proto-http==1.45.0 +opentelemetry-proto==1.45.0 +opentelemetry-sdk==1.45.0 +opentelemetry-semantic-conventions==0.66b0 +opt-einsum==3.4.0 +optax==0.2.4 +orbax-checkpoint==0.11.13 +orjson==3.12.0 +packaging==26.3 +pandas==2.2.3 +pillow==12.3.0 +platformdirs==4.11.14 +pluggy==1.6.0 +-e file:///home/gotham/tmp/plugrl/plugrl-protocol +-e file:///home/gotham/tmp/plugrl/plugrl-server +pretty-errors==1.2.25 +propcache==0.5.2 +protobuf==7.36.1 +pyarrow==20.0.0 +pydantic==2.13.5 +pydantic-core==2.46.5 +pygments==2.21.0 +pyparsing==3.3.3 +pysocks==1.7.1 +pytest==9.1.1 +python-dateutil==2.9.0.post0 +pytz==2026.3.post1 +pyvers==0.2.3 +pyyaml==6.0.3 +ray==2.58.0 +referencing==0.37.0 +regex==2026.9.10 +requests==2.34.2 +rich==15.0.0 +rpds-py==2026.6.3 +safetensors==0.5.3 +scipy==1.17.1 +sentencepiece==0.2.2 +setuptools==84.0.0 +simplejson==4.1.2 +six==1.17.0 +soupsieve==2.10 +sympy==1.14.0 +tensorboard==2.21.0 +tensorboard-data-server==0.7.2 +tensordict==0.14.2 +tensorstore==0.1.74 +tokenizers==0.21.4 +toml==0.10.2 +toolz==1.1.0 +torch==2.7.1 +torchvision==0.22.1 +tqdm==4.70.1 +tqdm-loggable==0.4.1 +transformers==4.53.2 +treescope==0.1.10 +triton==3.3.1 +typeguard==4.6.0 +typing-extensions==4.16.0 +typing-inspect==0.9.0 +typing-inspection==0.4.4 +tyro==1.0.16 +tzdata==2026.4 +urllib3==2.7.0 +wadler-lindig==0.1.7 +wandb==0.30.0 +websockets==17.1 +werkzeug==3.1.8 +wrapt==2.4.1 +xxhash==3.5.0 +yarl==1.24.5 +zipp==4.1.0 diff --git a/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it10.txt b/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it10.txt new file mode 100644 index 0000000..914f258 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it10.txt @@ -0,0 +1,109 @@ +56d0b8915d772035604fc2a55b057c3ed25b23e0d8e4f356c374231e240a4370 plugrl_server/algorithm/base_algorithm.py +2d59398047a868eb2aca50fe8a99c236f487807b9f514870ac4073644da65d86 plugrl_server/algorithm/distributed.py +4bb9d3f2b912935c4750e63e76791a317e88af7a2cbd93b9a653095b0a87c1db plugrl_server/algorithm/dppo/dppo_buffer.py +8ab1759a35534eee5ca524776e8ac53d2ee904f7e2ef710640261c9986922337 plugrl_server/algorithm/dppo/dppo_config.py +d77618cf493262c825a80ed428f375d77870ff67b776c748574ccca8a2d4af5a plugrl_server/algorithm/dppo/dppo_dist_config.py +65ab06f4e6151f5af9824654f49ce5267b8d06a0bed0221a37cd5cb1b2523318 plugrl_server/algorithm/dppo/dppo_dist.py +b280bcb5ad5ee3ffae9360a0f67819e032f9929f2bef81a02875870848059ff7 plugrl_server/algorithm/dppo/dppo_optimizer.py +e78705f16e2aa96df76eca3f75b650614395268cf7dc0e613b577f94e6705d85 plugrl_server/algorithm/dppo/dppo.py +1e46c7abebd22ac79fb8b228e79c30d623fca354d899ec83851fe3b58a3ad290 plugrl_server/algorithm/dppo/dppo_scheduler.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/dppo/__init__.py +7a545cd3d67dd85285a31a459f99db409fbd59a90104aefdb83727a6020ba63e plugrl_server/algorithm/dppo/third_party/__init__.py +7b3f14ce7e5d426a7814841113d0f3435e6f5733188159d4eb1118fc76ddcf00 plugrl_server/algorithm/dppo/third_party/reward_scaling.py +41b75540628b2486f9e0d451577b4fd575321956fc07067c6afe9eb2d3daa24f plugrl_server/algorithm/dppo/third_party/scheduler.py +587879b7141e202debac9013c483d9098c71e7067adec1a468e9276f04b58da1 plugrl_server/algorithm/dppo/train_state.py +b66c1acb2b366b5924ad30ea8f9e1e57da19c84a42b25a0d3df814eca8049036 plugrl_server/algorithm/dummy_algorithm.py +5f10ea70b4ec2011860fe845afe8e19b0ce7d9d9e5a7c3512f724d556ae7a901 plugrl_server/algorithm/evaluation.py +70af9ed7e56f18e932f89357cfc22ae84de2b4405159e5f0d741aff2c63e0de4 plugrl_server/algorithm/fpo/fpo_buffer.py +f1cc2b4f4df868588b2f321bd534790aaf92217a68f51a5cc4f99112b02c074b plugrl_server/algorithm/fpo/fpo_config.py +42aa760953489a83068232966c17161f70c44b5d3f058c8a961c28f13e15d79d plugrl_server/algorithm/fpo/fpo.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/fpo/__init__.py +ea71894ef655679feb7c82603c8465e9d1c304462f8b9ce3853afe990c85fea4 plugrl_server/algorithm/fpo/utils.py +818e44581ff29a748b4f9077733cdb78683c7a8dea2926383f3a3008b9112b31 plugrl_server/algorithm/__init__.py +794b1dd7b8e76b875552f69cdb9130e119871e8df01a60f3b046ba6ce968651d plugrl_server/algorithm/master_weights.py +989cc5493e0f631dbded20f55be99331a2b9daacd465514d3bcdce3c157d5347 plugrl_server/algorithm/registration.py +3cbcbf83312209b6b66e9bee3ac2097404185cb1e106062a105ca00fd4d3a891 plugrl_server/algorithm/train_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/buffer/__init__.py +a5f1f1838a30089df37c6716e16643736bd484adedd5f0014de6f76b0d6bc03e plugrl_server/buffer/numpy_tree_storage.py +46c71b28e589467742d7733e47512d7e66ff8003ded6ff9285b74e60481e147e plugrl_server/buffer/replay_buffer.py +f3eb2f00575f034517b1a0fac781f500bc7ce42bde21331cf8583c69bf9ce904 plugrl_server/buffer/rollout_buffer.py +534e3131a529999fd66bdb5213d893300b26856e8859154767bdefe1d2875983 plugrl_server/buffer/schema_migration.py +140a59a33260faf62b5d6cc04cc568935d7536129583057b01dc7740f71a737e plugrl_server/cli.py +64ff433c72a4e168fc6ee282e4b79277366ee63455a4e33fab82b11d0521a5da plugrl_server/cli_ray.py +57f2f5dd47ac6eb86d9264d3424972f0124deaa9246d05dfc2aff7926ba5bfca plugrl_server/common/checkpoint_manager.py +75bd8edd84b1c815905a1f6127b2e0a63429fe86d17b09f690ad372ec6f64856 plugrl_server/common/data_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/common/__init__.py +598e6f4f4593a63f61f31fcb451c7be0b00e21519fb3fdbcf5dee896833a8aeb plugrl_server/common/logging_utils.py +90e210352d7a19ec214c5f32b8e53d9b2becd8bd811b8eee9e1c0ab652d1fab5 plugrl_server/common/metrics.py +4d39436329efacdecbac9a6e61e5ecd535156abc80eb599a235e20e06b532ac8 plugrl_server/common/progress.py +fbf799d6070db6589990310ecf0ff16ad5947968438ec51f589f3c8342383003 plugrl_server/common/seeding.py +d0b10a81603604db0b9f96c3e42b12d8df55da7e0216cd25e5fefb3e27ad69a3 plugrl_server/common/tyro_utils.py +ee7db32bd24d367424f65dfe8cbe3a32e06fb1e72c13937041f1131059d79385 plugrl_server/__init__.py +fabb97e4177e40ea02e55087ea0dadf86845c175be35746663b2777251f9183f plugrl_server/paths.py +7b94db9740c293131ad78a2b90d66fb5e365d93a915a1c7397de7b491c8ec504 plugrl_server/policy/base_policy_gradient_diffusion_policy.py +0fb34bed9f82c677bc093c44b26fd663300b974aa26562ae2dc309b430b9ca85 plugrl_server/policy/base_policy_gradient_flow_policy.py +270cac6e5cb528ebd5aec9b32801976c3457495ae48c5f32a127cb0cbd2f3aba plugrl_server/policy/base_policy.py +adc80024a0b2c2e742de2cc36f0ab05f16c7d71931373a320eac85f152379ae2 plugrl_server/policy/base_torch_policy.py +66f0f7ed11cabfa4b6f0b554099f502e593db46e0cca886c9074d6b5604a0c62 plugrl_server/policy/dppo/dppo_policy.py +2e0a2d6b741ae2762fabcdf75e8f4569ef622d03f2360d89eb87e6bde2f997e2 plugrl_server/policy/dppo/__init__.py +63b4ebbd00f40654ad231ff1554d22485f72d5a13d5e755c9aec54175f6e62a0 plugrl_server/policy/dummy_policy.py +5042e45affe5e7e82ec71f16ca1476e4656351d0525f911fd91a0a65970b9b62 plugrl_server/policy/fpo/fpo_policy.py +66871189b4f92dc9b6b16f3eb38830c922f363a57a76dd464a0a6bbfe0f38750 plugrl_server/policy/fpo/__init__.py +a0fb32c5936e255b7e18af486cd82daaeccb829c655e6050ac369352ef4f77fd plugrl_server/policy/freezing.py +82ef5d06a4d1dfeca4e3cd5bebf1f6b69976d131d932288ea9eb1ee7833d4467 plugrl_server/policy/__init__.py +722bd543b81d3844ac30c37776711058f62d362177d0bb435cb2b29e11c2df9e plugrl_server/policy/openpi/__init__.py +5c4f99aaa3f29f60ad716d8f22105597824604e02bf03ff8cb2c361c02a14acb plugrl_server/policy/openpi/openpi_policy.py +6529cf6fac488fb78bb7730f55707c0cd1eb43ca6e91d7d7d43e36aec5d6fff8 plugrl_server/policy/openpi/openpi_transforming.py +82636d45ec176cc4a73886afe400d0c0f195a1f7850382be59c3ee1a4ca4e9db plugrl_server/policy/openpi/value_head.py +d42f35828ef2b6b9b940584bf1d843e8c8468a034dc885bc093815d4dca378dd plugrl_server/policy/registration.py +5e0ae56e52068e44dea0bf780bbd120046dc1e2b0f517f875fee0de02c80543b plugrl_server/policy/state_adapter.py +7d30e5453029d592fcd7c2d873731b72cf483ce99fe030a963f15ae1bb258332 plugrl_server/policy/state.py +a3e5175ce36c1eb9b3e384aa731d69d33d0863e22bdb9313c4a35d85838e212f plugrl_server/server/inference_coordinator.py +a3c4d0ab10053deacd07b3d32283c86ec30fe31fa682128467d0cb08e43ecd3c plugrl_server/server/lifecycle.py +503ccb909a4c3339486af819a1f3d6a56b3100e1280ebb089138b53c351c10aa plugrl_server/server/metadata.py +17de67f4ae7e7deb6b947c73f1077817f2a5ea6d45e4e18ba5a032b0fd102143 plugrl_server/server/protocol.py +d2d89fea8d4e4a10bb3df8628c46efb6380e021d01157fd96924686525158f4b plugrl_server/server/ray_agent_server.py +9d27aeec420c86dd72adb8c3f5951324ce53d46087644faf11cfc2d6c0400862 plugrl_server/server/ray_learner.py +e3fdc45689f05c39385462dda1e88d333a055af114bd752f0be9e79abb66fe97 plugrl_server/server/runtime_metrics.py +803a64912184047e79d40cc70192b022f6c656248f7acdc34c72d15fd0eb1d1e plugrl_server/server/runtime_scheduler.py +6076489eed897d003db73f07b959715987a1174666091b1a71e74bf41f0c14fe plugrl_server/server/training_backend.py +a322143fba9ea3336fc758f3c5250a2f0dc0a52e009d5b9c23148369113128ec plugrl_server/server/websocket_agent_server.py +a5ed9fc50e778b12fade66ac4e853d52cba96649ff5b4cc4b39d9e9b6e9e2430 plugrl_env_client/agent/base_agent.py +7e450fee0ecb6419855e72c1171bd3b7e4e06edade38c6bb5fa431d335338e91 plugrl_env_client/agent/websocket_env_client_agent.py +b3680448165e92d540867e77b2e96769bde9ea52f46845aa754986c3041d434a plugrl_env_client/cli.py +3648cb74799531826eea5440225b6cf4792c3e4a1244ae28c0a1743eea99725a plugrl_env_client/envs/atari/atari_env.py +bd257acc1c5c30f39f766f96dd4af58bfc13ae3b6ff0f012a421dbc45d707485 plugrl_env_client/envs/atari/atari_wrappers.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/atari/__init__.py +4132a1b69d36e9f8c28cd3a6d873b8af81b2722630a49ce294af4dbc47096123 plugrl_env_client/envs/base_env.py +5326f351d101bddebf15899c790f67fbb617ec8cd59d9d3bd7b0bdafdbfadee8 plugrl_env_client/envs/classic/classic_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/classic/__init__.py +616812fa3e47727ecf32d58193ea82d7491fa72e7248c14693657120c61f53e4 plugrl_env_client/envs/d4rl/d4rl_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/d4rl/__init__.py +e0461287f53363eabbb8503418e0de20c93279cdb865e6eb8b4016ba56bd51fd plugrl_env_client/envs/dummy_env.py +2162e3cf76b75ffe347adbf39dbd87e23eab36117cce8626d6145049683cbb9e plugrl_env_client/envs/__init__.py +4f308a45cc7c78c3617d172ccc60305efebba82605b6f32737ac75c8a6ea5ddc plugrl_env_client/envs/libero/image_tools.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/libero/__init__.py +f835ff5b01161df7c8c83ad76f06a7851c5306c2ccb1b35721db0a470552e0f8 plugrl_env_client/envs/libero/libero_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/mujoco/__init__.py +ea270ca84b88c4857ad02bcdf3dce7320057b85962f149d06a9261b6cfa0c862 plugrl_env_client/envs/mujoco/mujoco_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/robomimic/__init__.py +22c384593ac46df55c83812bfe9697c86c75e9a5cf487c09458bd5fbe120a353 plugrl_env_client/envs/robomimic/robomimic_env.py +8435d17ff7f893138a198c0452abbd1d98a51bb80a0cac362801964c280f6844 plugrl_env_client/__init__.py +1c07862ce6a46e8c25e271d0061f0b82333504d171e78a280188ea394f9308cb plugrl_env_client/recorder/args.py +d3c58231640d9bce38619c5c6d56e1d348177b6dbce5537f8a4fd10c382df883 plugrl_env_client/recorder/events.py +95b5fa86e4b13985b58e79bc1605fc91aed42418db67ba0e64b85cde0f93f943 plugrl_env_client/recorder/__init__.py +2f2a6a94ae4566a95bfece985572eddb51aeef8335bd8fe67fc5753417deae5c plugrl_env_client/recorder/recorder.py +b8d0b07d090eb0f97a50e7f7e542376b63cbd507dcd8ffb79fcaec1f1c744dc5 plugrl_env_client/recorder/sink.py +e3e42f2354b9f83eb15f99ad537f389cc29b4411fd173cf676c711037dea1176 plugrl_env_client/recorder/writers.py +e5d0ed0b9a141e762ff85dc96471a495c85bf9869084192518221c4c86ae61c6 plugrl_env_client/runner/args.py +fb6a4761fb3bd68440546ba1f4ace66ad4be1014ac0b5668041c06e3d2eac117 plugrl_env_client/runner/__init__.py +257b347e940daeefb025f3fa3ecca7316849a50d537bdd313473c86d692e998c plugrl_env_client/runner/rollout.py +66cbdf5d1572c987bfedd3ffad896aafb4b8f865dcbc6876e7e8037193367607 plugrl_env_client/runner/run.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/utils/__init__.py +76908794d381ef84755471be61ecc86bfe973eb42501541ce7815d786156df7f plugrl_env_client/utils/recorder.py +e75f1539456ac6b2054d0a3d7fa5886f672783b75975da111b86daacc5cb8bb6 plugrl_env_client/utils/registration.py +99c3587c966a918b83a65de63556fb58e24761055e07ecf32a7f269be268ec5a plugrl_env_client/utils/rollout.py +1a5766bb33b2e43472980c9ef7e76367c7ff06aeeafd4db1fe5cf134894772e7 plugrl_env_client/utils/wrappers/episode_stats_wrapper.py +65ebc060d3c5c1898ef54298e5e0e9588431ada249a90583180fbe1056fddb4e plugrl_env_client/utils/wrappers/__init__.py +e246f2ac99b771225302b87952a4050a7514e9baeb2dba5a1240af50f98ec240 plugrl_env_client/utils/wrappers/real_time_wrapper.py +827c8f87b7f56396885a1222d02d098185af675278b452dcc2124c49cbe1b9e9 plugrl_env_client/utils/wrappers/time_limit_wrapper.py diff --git a/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it5.txt b/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it5.txt new file mode 100644 index 0000000..914f258 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-source-e36-dppo-it5.txt @@ -0,0 +1,109 @@ +56d0b8915d772035604fc2a55b057c3ed25b23e0d8e4f356c374231e240a4370 plugrl_server/algorithm/base_algorithm.py +2d59398047a868eb2aca50fe8a99c236f487807b9f514870ac4073644da65d86 plugrl_server/algorithm/distributed.py +4bb9d3f2b912935c4750e63e76791a317e88af7a2cbd93b9a653095b0a87c1db plugrl_server/algorithm/dppo/dppo_buffer.py +8ab1759a35534eee5ca524776e8ac53d2ee904f7e2ef710640261c9986922337 plugrl_server/algorithm/dppo/dppo_config.py +d77618cf493262c825a80ed428f375d77870ff67b776c748574ccca8a2d4af5a plugrl_server/algorithm/dppo/dppo_dist_config.py +65ab06f4e6151f5af9824654f49ce5267b8d06a0bed0221a37cd5cb1b2523318 plugrl_server/algorithm/dppo/dppo_dist.py +b280bcb5ad5ee3ffae9360a0f67819e032f9929f2bef81a02875870848059ff7 plugrl_server/algorithm/dppo/dppo_optimizer.py +e78705f16e2aa96df76eca3f75b650614395268cf7dc0e613b577f94e6705d85 plugrl_server/algorithm/dppo/dppo.py +1e46c7abebd22ac79fb8b228e79c30d623fca354d899ec83851fe3b58a3ad290 plugrl_server/algorithm/dppo/dppo_scheduler.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/dppo/__init__.py +7a545cd3d67dd85285a31a459f99db409fbd59a90104aefdb83727a6020ba63e plugrl_server/algorithm/dppo/third_party/__init__.py +7b3f14ce7e5d426a7814841113d0f3435e6f5733188159d4eb1118fc76ddcf00 plugrl_server/algorithm/dppo/third_party/reward_scaling.py +41b75540628b2486f9e0d451577b4fd575321956fc07067c6afe9eb2d3daa24f plugrl_server/algorithm/dppo/third_party/scheduler.py +587879b7141e202debac9013c483d9098c71e7067adec1a468e9276f04b58da1 plugrl_server/algorithm/dppo/train_state.py +b66c1acb2b366b5924ad30ea8f9e1e57da19c84a42b25a0d3df814eca8049036 plugrl_server/algorithm/dummy_algorithm.py +5f10ea70b4ec2011860fe845afe8e19b0ce7d9d9e5a7c3512f724d556ae7a901 plugrl_server/algorithm/evaluation.py +70af9ed7e56f18e932f89357cfc22ae84de2b4405159e5f0d741aff2c63e0de4 plugrl_server/algorithm/fpo/fpo_buffer.py +f1cc2b4f4df868588b2f321bd534790aaf92217a68f51a5cc4f99112b02c074b plugrl_server/algorithm/fpo/fpo_config.py +42aa760953489a83068232966c17161f70c44b5d3f058c8a961c28f13e15d79d plugrl_server/algorithm/fpo/fpo.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/fpo/__init__.py +ea71894ef655679feb7c82603c8465e9d1c304462f8b9ce3853afe990c85fea4 plugrl_server/algorithm/fpo/utils.py +818e44581ff29a748b4f9077733cdb78683c7a8dea2926383f3a3008b9112b31 plugrl_server/algorithm/__init__.py +794b1dd7b8e76b875552f69cdb9130e119871e8df01a60f3b046ba6ce968651d plugrl_server/algorithm/master_weights.py +989cc5493e0f631dbded20f55be99331a2b9daacd465514d3bcdce3c157d5347 plugrl_server/algorithm/registration.py +3cbcbf83312209b6b66e9bee3ac2097404185cb1e106062a105ca00fd4d3a891 plugrl_server/algorithm/train_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/buffer/__init__.py +a5f1f1838a30089df37c6716e16643736bd484adedd5f0014de6f76b0d6bc03e plugrl_server/buffer/numpy_tree_storage.py +46c71b28e589467742d7733e47512d7e66ff8003ded6ff9285b74e60481e147e plugrl_server/buffer/replay_buffer.py +f3eb2f00575f034517b1a0fac781f500bc7ce42bde21331cf8583c69bf9ce904 plugrl_server/buffer/rollout_buffer.py +534e3131a529999fd66bdb5213d893300b26856e8859154767bdefe1d2875983 plugrl_server/buffer/schema_migration.py +140a59a33260faf62b5d6cc04cc568935d7536129583057b01dc7740f71a737e plugrl_server/cli.py +64ff433c72a4e168fc6ee282e4b79277366ee63455a4e33fab82b11d0521a5da plugrl_server/cli_ray.py +57f2f5dd47ac6eb86d9264d3424972f0124deaa9246d05dfc2aff7926ba5bfca plugrl_server/common/checkpoint_manager.py +75bd8edd84b1c815905a1f6127b2e0a63429fe86d17b09f690ad372ec6f64856 plugrl_server/common/data_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/common/__init__.py +598e6f4f4593a63f61f31fcb451c7be0b00e21519fb3fdbcf5dee896833a8aeb plugrl_server/common/logging_utils.py +90e210352d7a19ec214c5f32b8e53d9b2becd8bd811b8eee9e1c0ab652d1fab5 plugrl_server/common/metrics.py +4d39436329efacdecbac9a6e61e5ecd535156abc80eb599a235e20e06b532ac8 plugrl_server/common/progress.py +fbf799d6070db6589990310ecf0ff16ad5947968438ec51f589f3c8342383003 plugrl_server/common/seeding.py +d0b10a81603604db0b9f96c3e42b12d8df55da7e0216cd25e5fefb3e27ad69a3 plugrl_server/common/tyro_utils.py +ee7db32bd24d367424f65dfe8cbe3a32e06fb1e72c13937041f1131059d79385 plugrl_server/__init__.py +fabb97e4177e40ea02e55087ea0dadf86845c175be35746663b2777251f9183f plugrl_server/paths.py +7b94db9740c293131ad78a2b90d66fb5e365d93a915a1c7397de7b491c8ec504 plugrl_server/policy/base_policy_gradient_diffusion_policy.py +0fb34bed9f82c677bc093c44b26fd663300b974aa26562ae2dc309b430b9ca85 plugrl_server/policy/base_policy_gradient_flow_policy.py +270cac6e5cb528ebd5aec9b32801976c3457495ae48c5f32a127cb0cbd2f3aba plugrl_server/policy/base_policy.py +adc80024a0b2c2e742de2cc36f0ab05f16c7d71931373a320eac85f152379ae2 plugrl_server/policy/base_torch_policy.py +66f0f7ed11cabfa4b6f0b554099f502e593db46e0cca886c9074d6b5604a0c62 plugrl_server/policy/dppo/dppo_policy.py +2e0a2d6b741ae2762fabcdf75e8f4569ef622d03f2360d89eb87e6bde2f997e2 plugrl_server/policy/dppo/__init__.py +63b4ebbd00f40654ad231ff1554d22485f72d5a13d5e755c9aec54175f6e62a0 plugrl_server/policy/dummy_policy.py +5042e45affe5e7e82ec71f16ca1476e4656351d0525f911fd91a0a65970b9b62 plugrl_server/policy/fpo/fpo_policy.py +66871189b4f92dc9b6b16f3eb38830c922f363a57a76dd464a0a6bbfe0f38750 plugrl_server/policy/fpo/__init__.py +a0fb32c5936e255b7e18af486cd82daaeccb829c655e6050ac369352ef4f77fd plugrl_server/policy/freezing.py +82ef5d06a4d1dfeca4e3cd5bebf1f6b69976d131d932288ea9eb1ee7833d4467 plugrl_server/policy/__init__.py +722bd543b81d3844ac30c37776711058f62d362177d0bb435cb2b29e11c2df9e plugrl_server/policy/openpi/__init__.py +5c4f99aaa3f29f60ad716d8f22105597824604e02bf03ff8cb2c361c02a14acb plugrl_server/policy/openpi/openpi_policy.py +6529cf6fac488fb78bb7730f55707c0cd1eb43ca6e91d7d7d43e36aec5d6fff8 plugrl_server/policy/openpi/openpi_transforming.py +82636d45ec176cc4a73886afe400d0c0f195a1f7850382be59c3ee1a4ca4e9db plugrl_server/policy/openpi/value_head.py +d42f35828ef2b6b9b940584bf1d843e8c8468a034dc885bc093815d4dca378dd plugrl_server/policy/registration.py +5e0ae56e52068e44dea0bf780bbd120046dc1e2b0f517f875fee0de02c80543b plugrl_server/policy/state_adapter.py +7d30e5453029d592fcd7c2d873731b72cf483ce99fe030a963f15ae1bb258332 plugrl_server/policy/state.py +a3e5175ce36c1eb9b3e384aa731d69d33d0863e22bdb9313c4a35d85838e212f plugrl_server/server/inference_coordinator.py +a3c4d0ab10053deacd07b3d32283c86ec30fe31fa682128467d0cb08e43ecd3c plugrl_server/server/lifecycle.py +503ccb909a4c3339486af819a1f3d6a56b3100e1280ebb089138b53c351c10aa plugrl_server/server/metadata.py +17de67f4ae7e7deb6b947c73f1077817f2a5ea6d45e4e18ba5a032b0fd102143 plugrl_server/server/protocol.py +d2d89fea8d4e4a10bb3df8628c46efb6380e021d01157fd96924686525158f4b plugrl_server/server/ray_agent_server.py +9d27aeec420c86dd72adb8c3f5951324ce53d46087644faf11cfc2d6c0400862 plugrl_server/server/ray_learner.py +e3fdc45689f05c39385462dda1e88d333a055af114bd752f0be9e79abb66fe97 plugrl_server/server/runtime_metrics.py +803a64912184047e79d40cc70192b022f6c656248f7acdc34c72d15fd0eb1d1e plugrl_server/server/runtime_scheduler.py +6076489eed897d003db73f07b959715987a1174666091b1a71e74bf41f0c14fe plugrl_server/server/training_backend.py +a322143fba9ea3336fc758f3c5250a2f0dc0a52e009d5b9c23148369113128ec plugrl_server/server/websocket_agent_server.py +a5ed9fc50e778b12fade66ac4e853d52cba96649ff5b4cc4b39d9e9b6e9e2430 plugrl_env_client/agent/base_agent.py +7e450fee0ecb6419855e72c1171bd3b7e4e06edade38c6bb5fa431d335338e91 plugrl_env_client/agent/websocket_env_client_agent.py +b3680448165e92d540867e77b2e96769bde9ea52f46845aa754986c3041d434a plugrl_env_client/cli.py +3648cb74799531826eea5440225b6cf4792c3e4a1244ae28c0a1743eea99725a plugrl_env_client/envs/atari/atari_env.py +bd257acc1c5c30f39f766f96dd4af58bfc13ae3b6ff0f012a421dbc45d707485 plugrl_env_client/envs/atari/atari_wrappers.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/atari/__init__.py +4132a1b69d36e9f8c28cd3a6d873b8af81b2722630a49ce294af4dbc47096123 plugrl_env_client/envs/base_env.py +5326f351d101bddebf15899c790f67fbb617ec8cd59d9d3bd7b0bdafdbfadee8 plugrl_env_client/envs/classic/classic_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/classic/__init__.py +616812fa3e47727ecf32d58193ea82d7491fa72e7248c14693657120c61f53e4 plugrl_env_client/envs/d4rl/d4rl_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/d4rl/__init__.py +e0461287f53363eabbb8503418e0de20c93279cdb865e6eb8b4016ba56bd51fd plugrl_env_client/envs/dummy_env.py +2162e3cf76b75ffe347adbf39dbd87e23eab36117cce8626d6145049683cbb9e plugrl_env_client/envs/__init__.py +4f308a45cc7c78c3617d172ccc60305efebba82605b6f32737ac75c8a6ea5ddc plugrl_env_client/envs/libero/image_tools.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/libero/__init__.py +f835ff5b01161df7c8c83ad76f06a7851c5306c2ccb1b35721db0a470552e0f8 plugrl_env_client/envs/libero/libero_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/mujoco/__init__.py +ea270ca84b88c4857ad02bcdf3dce7320057b85962f149d06a9261b6cfa0c862 plugrl_env_client/envs/mujoco/mujoco_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/robomimic/__init__.py +22c384593ac46df55c83812bfe9697c86c75e9a5cf487c09458bd5fbe120a353 plugrl_env_client/envs/robomimic/robomimic_env.py +8435d17ff7f893138a198c0452abbd1d98a51bb80a0cac362801964c280f6844 plugrl_env_client/__init__.py +1c07862ce6a46e8c25e271d0061f0b82333504d171e78a280188ea394f9308cb plugrl_env_client/recorder/args.py +d3c58231640d9bce38619c5c6d56e1d348177b6dbce5537f8a4fd10c382df883 plugrl_env_client/recorder/events.py +95b5fa86e4b13985b58e79bc1605fc91aed42418db67ba0e64b85cde0f93f943 plugrl_env_client/recorder/__init__.py +2f2a6a94ae4566a95bfece985572eddb51aeef8335bd8fe67fc5753417deae5c plugrl_env_client/recorder/recorder.py +b8d0b07d090eb0f97a50e7f7e542376b63cbd507dcd8ffb79fcaec1f1c744dc5 plugrl_env_client/recorder/sink.py +e3e42f2354b9f83eb15f99ad537f389cc29b4411fd173cf676c711037dea1176 plugrl_env_client/recorder/writers.py +e5d0ed0b9a141e762ff85dc96471a495c85bf9869084192518221c4c86ae61c6 plugrl_env_client/runner/args.py +fb6a4761fb3bd68440546ba1f4ace66ad4be1014ac0b5668041c06e3d2eac117 plugrl_env_client/runner/__init__.py +257b347e940daeefb025f3fa3ecca7316849a50d537bdd313473c86d692e998c plugrl_env_client/runner/rollout.py +66cbdf5d1572c987bfedd3ffad896aafb4b8f865dcbc6876e7e8037193367607 plugrl_env_client/runner/run.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/utils/__init__.py +76908794d381ef84755471be61ecc86bfe973eb42501541ce7815d786156df7f plugrl_env_client/utils/recorder.py +e75f1539456ac6b2054d0a3d7fa5886f672783b75975da111b86daacc5cb8bb6 plugrl_env_client/utils/registration.py +99c3587c966a918b83a65de63556fb58e24761055e07ecf32a7f269be268ec5a plugrl_env_client/utils/rollout.py +1a5766bb33b2e43472980c9ef7e76367c7ff06aeeafd4db1fe5cf134894772e7 plugrl_env_client/utils/wrappers/episode_stats_wrapper.py +65ebc060d3c5c1898ef54298e5e0e9588431ada249a90583180fbe1056fddb4e plugrl_env_client/utils/wrappers/__init__.py +e246f2ac99b771225302b87952a4050a7514e9baeb2dba5a1240af50f98ec240 plugrl_env_client/utils/wrappers/real_time_wrapper.py +827c8f87b7f56396885a1222d02d098185af675278b452dcc2124c49cbe1b9e9 plugrl_env_client/utils/wrappers/time_limit_wrapper.py diff --git a/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s7-it5.txt b/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s7-it5.txt new file mode 100644 index 0000000..914f258 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s7-it5.txt @@ -0,0 +1,109 @@ +56d0b8915d772035604fc2a55b057c3ed25b23e0d8e4f356c374231e240a4370 plugrl_server/algorithm/base_algorithm.py +2d59398047a868eb2aca50fe8a99c236f487807b9f514870ac4073644da65d86 plugrl_server/algorithm/distributed.py +4bb9d3f2b912935c4750e63e76791a317e88af7a2cbd93b9a653095b0a87c1db plugrl_server/algorithm/dppo/dppo_buffer.py +8ab1759a35534eee5ca524776e8ac53d2ee904f7e2ef710640261c9986922337 plugrl_server/algorithm/dppo/dppo_config.py +d77618cf493262c825a80ed428f375d77870ff67b776c748574ccca8a2d4af5a plugrl_server/algorithm/dppo/dppo_dist_config.py +65ab06f4e6151f5af9824654f49ce5267b8d06a0bed0221a37cd5cb1b2523318 plugrl_server/algorithm/dppo/dppo_dist.py +b280bcb5ad5ee3ffae9360a0f67819e032f9929f2bef81a02875870848059ff7 plugrl_server/algorithm/dppo/dppo_optimizer.py +e78705f16e2aa96df76eca3f75b650614395268cf7dc0e613b577f94e6705d85 plugrl_server/algorithm/dppo/dppo.py +1e46c7abebd22ac79fb8b228e79c30d623fca354d899ec83851fe3b58a3ad290 plugrl_server/algorithm/dppo/dppo_scheduler.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/dppo/__init__.py +7a545cd3d67dd85285a31a459f99db409fbd59a90104aefdb83727a6020ba63e plugrl_server/algorithm/dppo/third_party/__init__.py +7b3f14ce7e5d426a7814841113d0f3435e6f5733188159d4eb1118fc76ddcf00 plugrl_server/algorithm/dppo/third_party/reward_scaling.py +41b75540628b2486f9e0d451577b4fd575321956fc07067c6afe9eb2d3daa24f plugrl_server/algorithm/dppo/third_party/scheduler.py +587879b7141e202debac9013c483d9098c71e7067adec1a468e9276f04b58da1 plugrl_server/algorithm/dppo/train_state.py +b66c1acb2b366b5924ad30ea8f9e1e57da19c84a42b25a0d3df814eca8049036 plugrl_server/algorithm/dummy_algorithm.py +5f10ea70b4ec2011860fe845afe8e19b0ce7d9d9e5a7c3512f724d556ae7a901 plugrl_server/algorithm/evaluation.py +70af9ed7e56f18e932f89357cfc22ae84de2b4405159e5f0d741aff2c63e0de4 plugrl_server/algorithm/fpo/fpo_buffer.py +f1cc2b4f4df868588b2f321bd534790aaf92217a68f51a5cc4f99112b02c074b plugrl_server/algorithm/fpo/fpo_config.py +42aa760953489a83068232966c17161f70c44b5d3f058c8a961c28f13e15d79d plugrl_server/algorithm/fpo/fpo.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/fpo/__init__.py +ea71894ef655679feb7c82603c8465e9d1c304462f8b9ce3853afe990c85fea4 plugrl_server/algorithm/fpo/utils.py +818e44581ff29a748b4f9077733cdb78683c7a8dea2926383f3a3008b9112b31 plugrl_server/algorithm/__init__.py +794b1dd7b8e76b875552f69cdb9130e119871e8df01a60f3b046ba6ce968651d plugrl_server/algorithm/master_weights.py +989cc5493e0f631dbded20f55be99331a2b9daacd465514d3bcdce3c157d5347 plugrl_server/algorithm/registration.py +3cbcbf83312209b6b66e9bee3ac2097404185cb1e106062a105ca00fd4d3a891 plugrl_server/algorithm/train_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/buffer/__init__.py +a5f1f1838a30089df37c6716e16643736bd484adedd5f0014de6f76b0d6bc03e plugrl_server/buffer/numpy_tree_storage.py +46c71b28e589467742d7733e47512d7e66ff8003ded6ff9285b74e60481e147e plugrl_server/buffer/replay_buffer.py +f3eb2f00575f034517b1a0fac781f500bc7ce42bde21331cf8583c69bf9ce904 plugrl_server/buffer/rollout_buffer.py +534e3131a529999fd66bdb5213d893300b26856e8859154767bdefe1d2875983 plugrl_server/buffer/schema_migration.py +140a59a33260faf62b5d6cc04cc568935d7536129583057b01dc7740f71a737e plugrl_server/cli.py +64ff433c72a4e168fc6ee282e4b79277366ee63455a4e33fab82b11d0521a5da plugrl_server/cli_ray.py +57f2f5dd47ac6eb86d9264d3424972f0124deaa9246d05dfc2aff7926ba5bfca plugrl_server/common/checkpoint_manager.py +75bd8edd84b1c815905a1f6127b2e0a63429fe86d17b09f690ad372ec6f64856 plugrl_server/common/data_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/common/__init__.py +598e6f4f4593a63f61f31fcb451c7be0b00e21519fb3fdbcf5dee896833a8aeb plugrl_server/common/logging_utils.py +90e210352d7a19ec214c5f32b8e53d9b2becd8bd811b8eee9e1c0ab652d1fab5 plugrl_server/common/metrics.py +4d39436329efacdecbac9a6e61e5ecd535156abc80eb599a235e20e06b532ac8 plugrl_server/common/progress.py +fbf799d6070db6589990310ecf0ff16ad5947968438ec51f589f3c8342383003 plugrl_server/common/seeding.py +d0b10a81603604db0b9f96c3e42b12d8df55da7e0216cd25e5fefb3e27ad69a3 plugrl_server/common/tyro_utils.py +ee7db32bd24d367424f65dfe8cbe3a32e06fb1e72c13937041f1131059d79385 plugrl_server/__init__.py +fabb97e4177e40ea02e55087ea0dadf86845c175be35746663b2777251f9183f plugrl_server/paths.py +7b94db9740c293131ad78a2b90d66fb5e365d93a915a1c7397de7b491c8ec504 plugrl_server/policy/base_policy_gradient_diffusion_policy.py +0fb34bed9f82c677bc093c44b26fd663300b974aa26562ae2dc309b430b9ca85 plugrl_server/policy/base_policy_gradient_flow_policy.py +270cac6e5cb528ebd5aec9b32801976c3457495ae48c5f32a127cb0cbd2f3aba plugrl_server/policy/base_policy.py +adc80024a0b2c2e742de2cc36f0ab05f16c7d71931373a320eac85f152379ae2 plugrl_server/policy/base_torch_policy.py +66f0f7ed11cabfa4b6f0b554099f502e593db46e0cca886c9074d6b5604a0c62 plugrl_server/policy/dppo/dppo_policy.py +2e0a2d6b741ae2762fabcdf75e8f4569ef622d03f2360d89eb87e6bde2f997e2 plugrl_server/policy/dppo/__init__.py +63b4ebbd00f40654ad231ff1554d22485f72d5a13d5e755c9aec54175f6e62a0 plugrl_server/policy/dummy_policy.py +5042e45affe5e7e82ec71f16ca1476e4656351d0525f911fd91a0a65970b9b62 plugrl_server/policy/fpo/fpo_policy.py +66871189b4f92dc9b6b16f3eb38830c922f363a57a76dd464a0a6bbfe0f38750 plugrl_server/policy/fpo/__init__.py +a0fb32c5936e255b7e18af486cd82daaeccb829c655e6050ac369352ef4f77fd plugrl_server/policy/freezing.py +82ef5d06a4d1dfeca4e3cd5bebf1f6b69976d131d932288ea9eb1ee7833d4467 plugrl_server/policy/__init__.py +722bd543b81d3844ac30c37776711058f62d362177d0bb435cb2b29e11c2df9e plugrl_server/policy/openpi/__init__.py +5c4f99aaa3f29f60ad716d8f22105597824604e02bf03ff8cb2c361c02a14acb plugrl_server/policy/openpi/openpi_policy.py +6529cf6fac488fb78bb7730f55707c0cd1eb43ca6e91d7d7d43e36aec5d6fff8 plugrl_server/policy/openpi/openpi_transforming.py +82636d45ec176cc4a73886afe400d0c0f195a1f7850382be59c3ee1a4ca4e9db plugrl_server/policy/openpi/value_head.py +d42f35828ef2b6b9b940584bf1d843e8c8468a034dc885bc093815d4dca378dd plugrl_server/policy/registration.py +5e0ae56e52068e44dea0bf780bbd120046dc1e2b0f517f875fee0de02c80543b plugrl_server/policy/state_adapter.py +7d30e5453029d592fcd7c2d873731b72cf483ce99fe030a963f15ae1bb258332 plugrl_server/policy/state.py +a3e5175ce36c1eb9b3e384aa731d69d33d0863e22bdb9313c4a35d85838e212f plugrl_server/server/inference_coordinator.py +a3c4d0ab10053deacd07b3d32283c86ec30fe31fa682128467d0cb08e43ecd3c plugrl_server/server/lifecycle.py +503ccb909a4c3339486af819a1f3d6a56b3100e1280ebb089138b53c351c10aa plugrl_server/server/metadata.py +17de67f4ae7e7deb6b947c73f1077817f2a5ea6d45e4e18ba5a032b0fd102143 plugrl_server/server/protocol.py +d2d89fea8d4e4a10bb3df8628c46efb6380e021d01157fd96924686525158f4b plugrl_server/server/ray_agent_server.py +9d27aeec420c86dd72adb8c3f5951324ce53d46087644faf11cfc2d6c0400862 plugrl_server/server/ray_learner.py +e3fdc45689f05c39385462dda1e88d333a055af114bd752f0be9e79abb66fe97 plugrl_server/server/runtime_metrics.py +803a64912184047e79d40cc70192b022f6c656248f7acdc34c72d15fd0eb1d1e plugrl_server/server/runtime_scheduler.py +6076489eed897d003db73f07b959715987a1174666091b1a71e74bf41f0c14fe plugrl_server/server/training_backend.py +a322143fba9ea3336fc758f3c5250a2f0dc0a52e009d5b9c23148369113128ec plugrl_server/server/websocket_agent_server.py +a5ed9fc50e778b12fade66ac4e853d52cba96649ff5b4cc4b39d9e9b6e9e2430 plugrl_env_client/agent/base_agent.py +7e450fee0ecb6419855e72c1171bd3b7e4e06edade38c6bb5fa431d335338e91 plugrl_env_client/agent/websocket_env_client_agent.py +b3680448165e92d540867e77b2e96769bde9ea52f46845aa754986c3041d434a plugrl_env_client/cli.py +3648cb74799531826eea5440225b6cf4792c3e4a1244ae28c0a1743eea99725a plugrl_env_client/envs/atari/atari_env.py +bd257acc1c5c30f39f766f96dd4af58bfc13ae3b6ff0f012a421dbc45d707485 plugrl_env_client/envs/atari/atari_wrappers.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/atari/__init__.py +4132a1b69d36e9f8c28cd3a6d873b8af81b2722630a49ce294af4dbc47096123 plugrl_env_client/envs/base_env.py +5326f351d101bddebf15899c790f67fbb617ec8cd59d9d3bd7b0bdafdbfadee8 plugrl_env_client/envs/classic/classic_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/classic/__init__.py +616812fa3e47727ecf32d58193ea82d7491fa72e7248c14693657120c61f53e4 plugrl_env_client/envs/d4rl/d4rl_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/d4rl/__init__.py +e0461287f53363eabbb8503418e0de20c93279cdb865e6eb8b4016ba56bd51fd plugrl_env_client/envs/dummy_env.py +2162e3cf76b75ffe347adbf39dbd87e23eab36117cce8626d6145049683cbb9e plugrl_env_client/envs/__init__.py +4f308a45cc7c78c3617d172ccc60305efebba82605b6f32737ac75c8a6ea5ddc plugrl_env_client/envs/libero/image_tools.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/libero/__init__.py +f835ff5b01161df7c8c83ad76f06a7851c5306c2ccb1b35721db0a470552e0f8 plugrl_env_client/envs/libero/libero_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/mujoco/__init__.py +ea270ca84b88c4857ad02bcdf3dce7320057b85962f149d06a9261b6cfa0c862 plugrl_env_client/envs/mujoco/mujoco_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/robomimic/__init__.py +22c384593ac46df55c83812bfe9697c86c75e9a5cf487c09458bd5fbe120a353 plugrl_env_client/envs/robomimic/robomimic_env.py +8435d17ff7f893138a198c0452abbd1d98a51bb80a0cac362801964c280f6844 plugrl_env_client/__init__.py +1c07862ce6a46e8c25e271d0061f0b82333504d171e78a280188ea394f9308cb plugrl_env_client/recorder/args.py +d3c58231640d9bce38619c5c6d56e1d348177b6dbce5537f8a4fd10c382df883 plugrl_env_client/recorder/events.py +95b5fa86e4b13985b58e79bc1605fc91aed42418db67ba0e64b85cde0f93f943 plugrl_env_client/recorder/__init__.py +2f2a6a94ae4566a95bfece985572eddb51aeef8335bd8fe67fc5753417deae5c plugrl_env_client/recorder/recorder.py +b8d0b07d090eb0f97a50e7f7e542376b63cbd507dcd8ffb79fcaec1f1c744dc5 plugrl_env_client/recorder/sink.py +e3e42f2354b9f83eb15f99ad537f389cc29b4411fd173cf676c711037dea1176 plugrl_env_client/recorder/writers.py +e5d0ed0b9a141e762ff85dc96471a495c85bf9869084192518221c4c86ae61c6 plugrl_env_client/runner/args.py +fb6a4761fb3bd68440546ba1f4ace66ad4be1014ac0b5668041c06e3d2eac117 plugrl_env_client/runner/__init__.py +257b347e940daeefb025f3fa3ecca7316849a50d537bdd313473c86d692e998c plugrl_env_client/runner/rollout.py +66cbdf5d1572c987bfedd3ffad896aafb4b8f865dcbc6876e7e8037193367607 plugrl_env_client/runner/run.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/utils/__init__.py +76908794d381ef84755471be61ecc86bfe973eb42501541ce7815d786156df7f plugrl_env_client/utils/recorder.py +e75f1539456ac6b2054d0a3d7fa5886f672783b75975da111b86daacc5cb8bb6 plugrl_env_client/utils/registration.py +99c3587c966a918b83a65de63556fb58e24761055e07ecf32a7f269be268ec5a plugrl_env_client/utils/rollout.py +1a5766bb33b2e43472980c9ef7e76367c7ff06aeeafd4db1fe5cf134894772e7 plugrl_env_client/utils/wrappers/episode_stats_wrapper.py +65ebc060d3c5c1898ef54298e5e0e9588431ada249a90583180fbe1056fddb4e plugrl_env_client/utils/wrappers/__init__.py +e246f2ac99b771225302b87952a4050a7514e9baeb2dba5a1240af50f98ec240 plugrl_env_client/utils/wrappers/real_time_wrapper.py +827c8f87b7f56396885a1222d02d098185af675278b452dcc2124c49cbe1b9e9 plugrl_env_client/utils/wrappers/time_limit_wrapper.py diff --git a/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s8-it5.txt b/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s8-it5.txt new file mode 100644 index 0000000..914f258 --- /dev/null +++ b/experiments/e36-pi0-longer/results/environment-source-e36-fpopp-s8-it5.txt @@ -0,0 +1,109 @@ +56d0b8915d772035604fc2a55b057c3ed25b23e0d8e4f356c374231e240a4370 plugrl_server/algorithm/base_algorithm.py +2d59398047a868eb2aca50fe8a99c236f487807b9f514870ac4073644da65d86 plugrl_server/algorithm/distributed.py +4bb9d3f2b912935c4750e63e76791a317e88af7a2cbd93b9a653095b0a87c1db plugrl_server/algorithm/dppo/dppo_buffer.py +8ab1759a35534eee5ca524776e8ac53d2ee904f7e2ef710640261c9986922337 plugrl_server/algorithm/dppo/dppo_config.py +d77618cf493262c825a80ed428f375d77870ff67b776c748574ccca8a2d4af5a plugrl_server/algorithm/dppo/dppo_dist_config.py +65ab06f4e6151f5af9824654f49ce5267b8d06a0bed0221a37cd5cb1b2523318 plugrl_server/algorithm/dppo/dppo_dist.py +b280bcb5ad5ee3ffae9360a0f67819e032f9929f2bef81a02875870848059ff7 plugrl_server/algorithm/dppo/dppo_optimizer.py +e78705f16e2aa96df76eca3f75b650614395268cf7dc0e613b577f94e6705d85 plugrl_server/algorithm/dppo/dppo.py +1e46c7abebd22ac79fb8b228e79c30d623fca354d899ec83851fe3b58a3ad290 plugrl_server/algorithm/dppo/dppo_scheduler.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/dppo/__init__.py +7a545cd3d67dd85285a31a459f99db409fbd59a90104aefdb83727a6020ba63e plugrl_server/algorithm/dppo/third_party/__init__.py +7b3f14ce7e5d426a7814841113d0f3435e6f5733188159d4eb1118fc76ddcf00 plugrl_server/algorithm/dppo/third_party/reward_scaling.py +41b75540628b2486f9e0d451577b4fd575321956fc07067c6afe9eb2d3daa24f plugrl_server/algorithm/dppo/third_party/scheduler.py +587879b7141e202debac9013c483d9098c71e7067adec1a468e9276f04b58da1 plugrl_server/algorithm/dppo/train_state.py +b66c1acb2b366b5924ad30ea8f9e1e57da19c84a42b25a0d3df814eca8049036 plugrl_server/algorithm/dummy_algorithm.py +5f10ea70b4ec2011860fe845afe8e19b0ce7d9d9e5a7c3512f724d556ae7a901 plugrl_server/algorithm/evaluation.py +70af9ed7e56f18e932f89357cfc22ae84de2b4405159e5f0d741aff2c63e0de4 plugrl_server/algorithm/fpo/fpo_buffer.py +f1cc2b4f4df868588b2f321bd534790aaf92217a68f51a5cc4f99112b02c074b plugrl_server/algorithm/fpo/fpo_config.py +42aa760953489a83068232966c17161f70c44b5d3f058c8a961c28f13e15d79d plugrl_server/algorithm/fpo/fpo.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/algorithm/fpo/__init__.py +ea71894ef655679feb7c82603c8465e9d1c304462f8b9ce3853afe990c85fea4 plugrl_server/algorithm/fpo/utils.py +818e44581ff29a748b4f9077733cdb78683c7a8dea2926383f3a3008b9112b31 plugrl_server/algorithm/__init__.py +794b1dd7b8e76b875552f69cdb9130e119871e8df01a60f3b046ba6ce968651d plugrl_server/algorithm/master_weights.py +989cc5493e0f631dbded20f55be99331a2b9daacd465514d3bcdce3c157d5347 plugrl_server/algorithm/registration.py +3cbcbf83312209b6b66e9bee3ac2097404185cb1e106062a105ca00fd4d3a891 plugrl_server/algorithm/train_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/buffer/__init__.py +a5f1f1838a30089df37c6716e16643736bd484adedd5f0014de6f76b0d6bc03e plugrl_server/buffer/numpy_tree_storage.py +46c71b28e589467742d7733e47512d7e66ff8003ded6ff9285b74e60481e147e plugrl_server/buffer/replay_buffer.py +f3eb2f00575f034517b1a0fac781f500bc7ce42bde21331cf8583c69bf9ce904 plugrl_server/buffer/rollout_buffer.py +534e3131a529999fd66bdb5213d893300b26856e8859154767bdefe1d2875983 plugrl_server/buffer/schema_migration.py +140a59a33260faf62b5d6cc04cc568935d7536129583057b01dc7740f71a737e plugrl_server/cli.py +64ff433c72a4e168fc6ee282e4b79277366ee63455a4e33fab82b11d0521a5da plugrl_server/cli_ray.py +57f2f5dd47ac6eb86d9264d3424972f0124deaa9246d05dfc2aff7926ba5bfca plugrl_server/common/checkpoint_manager.py +75bd8edd84b1c815905a1f6127b2e0a63429fe86d17b09f690ad372ec6f64856 plugrl_server/common/data_utils.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_server/common/__init__.py +598e6f4f4593a63f61f31fcb451c7be0b00e21519fb3fdbcf5dee896833a8aeb plugrl_server/common/logging_utils.py +90e210352d7a19ec214c5f32b8e53d9b2becd8bd811b8eee9e1c0ab652d1fab5 plugrl_server/common/metrics.py +4d39436329efacdecbac9a6e61e5ecd535156abc80eb599a235e20e06b532ac8 plugrl_server/common/progress.py +fbf799d6070db6589990310ecf0ff16ad5947968438ec51f589f3c8342383003 plugrl_server/common/seeding.py +d0b10a81603604db0b9f96c3e42b12d8df55da7e0216cd25e5fefb3e27ad69a3 plugrl_server/common/tyro_utils.py +ee7db32bd24d367424f65dfe8cbe3a32e06fb1e72c13937041f1131059d79385 plugrl_server/__init__.py +fabb97e4177e40ea02e55087ea0dadf86845c175be35746663b2777251f9183f plugrl_server/paths.py +7b94db9740c293131ad78a2b90d66fb5e365d93a915a1c7397de7b491c8ec504 plugrl_server/policy/base_policy_gradient_diffusion_policy.py +0fb34bed9f82c677bc093c44b26fd663300b974aa26562ae2dc309b430b9ca85 plugrl_server/policy/base_policy_gradient_flow_policy.py +270cac6e5cb528ebd5aec9b32801976c3457495ae48c5f32a127cb0cbd2f3aba plugrl_server/policy/base_policy.py +adc80024a0b2c2e742de2cc36f0ab05f16c7d71931373a320eac85f152379ae2 plugrl_server/policy/base_torch_policy.py +66f0f7ed11cabfa4b6f0b554099f502e593db46e0cca886c9074d6b5604a0c62 plugrl_server/policy/dppo/dppo_policy.py +2e0a2d6b741ae2762fabcdf75e8f4569ef622d03f2360d89eb87e6bde2f997e2 plugrl_server/policy/dppo/__init__.py +63b4ebbd00f40654ad231ff1554d22485f72d5a13d5e755c9aec54175f6e62a0 plugrl_server/policy/dummy_policy.py +5042e45affe5e7e82ec71f16ca1476e4656351d0525f911fd91a0a65970b9b62 plugrl_server/policy/fpo/fpo_policy.py +66871189b4f92dc9b6b16f3eb38830c922f363a57a76dd464a0a6bbfe0f38750 plugrl_server/policy/fpo/__init__.py +a0fb32c5936e255b7e18af486cd82daaeccb829c655e6050ac369352ef4f77fd plugrl_server/policy/freezing.py +82ef5d06a4d1dfeca4e3cd5bebf1f6b69976d131d932288ea9eb1ee7833d4467 plugrl_server/policy/__init__.py +722bd543b81d3844ac30c37776711058f62d362177d0bb435cb2b29e11c2df9e plugrl_server/policy/openpi/__init__.py +5c4f99aaa3f29f60ad716d8f22105597824604e02bf03ff8cb2c361c02a14acb plugrl_server/policy/openpi/openpi_policy.py +6529cf6fac488fb78bb7730f55707c0cd1eb43ca6e91d7d7d43e36aec5d6fff8 plugrl_server/policy/openpi/openpi_transforming.py +82636d45ec176cc4a73886afe400d0c0f195a1f7850382be59c3ee1a4ca4e9db plugrl_server/policy/openpi/value_head.py +d42f35828ef2b6b9b940584bf1d843e8c8468a034dc885bc093815d4dca378dd plugrl_server/policy/registration.py +5e0ae56e52068e44dea0bf780bbd120046dc1e2b0f517f875fee0de02c80543b plugrl_server/policy/state_adapter.py +7d30e5453029d592fcd7c2d873731b72cf483ce99fe030a963f15ae1bb258332 plugrl_server/policy/state.py +a3e5175ce36c1eb9b3e384aa731d69d33d0863e22bdb9313c4a35d85838e212f plugrl_server/server/inference_coordinator.py +a3c4d0ab10053deacd07b3d32283c86ec30fe31fa682128467d0cb08e43ecd3c plugrl_server/server/lifecycle.py +503ccb909a4c3339486af819a1f3d6a56b3100e1280ebb089138b53c351c10aa plugrl_server/server/metadata.py +17de67f4ae7e7deb6b947c73f1077817f2a5ea6d45e4e18ba5a032b0fd102143 plugrl_server/server/protocol.py +d2d89fea8d4e4a10bb3df8628c46efb6380e021d01157fd96924686525158f4b plugrl_server/server/ray_agent_server.py +9d27aeec420c86dd72adb8c3f5951324ce53d46087644faf11cfc2d6c0400862 plugrl_server/server/ray_learner.py +e3fdc45689f05c39385462dda1e88d333a055af114bd752f0be9e79abb66fe97 plugrl_server/server/runtime_metrics.py +803a64912184047e79d40cc70192b022f6c656248f7acdc34c72d15fd0eb1d1e plugrl_server/server/runtime_scheduler.py +6076489eed897d003db73f07b959715987a1174666091b1a71e74bf41f0c14fe plugrl_server/server/training_backend.py +a322143fba9ea3336fc758f3c5250a2f0dc0a52e009d5b9c23148369113128ec plugrl_server/server/websocket_agent_server.py +a5ed9fc50e778b12fade66ac4e853d52cba96649ff5b4cc4b39d9e9b6e9e2430 plugrl_env_client/agent/base_agent.py +7e450fee0ecb6419855e72c1171bd3b7e4e06edade38c6bb5fa431d335338e91 plugrl_env_client/agent/websocket_env_client_agent.py +b3680448165e92d540867e77b2e96769bde9ea52f46845aa754986c3041d434a plugrl_env_client/cli.py +3648cb74799531826eea5440225b6cf4792c3e4a1244ae28c0a1743eea99725a plugrl_env_client/envs/atari/atari_env.py +bd257acc1c5c30f39f766f96dd4af58bfc13ae3b6ff0f012a421dbc45d707485 plugrl_env_client/envs/atari/atari_wrappers.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/atari/__init__.py +4132a1b69d36e9f8c28cd3a6d873b8af81b2722630a49ce294af4dbc47096123 plugrl_env_client/envs/base_env.py +5326f351d101bddebf15899c790f67fbb617ec8cd59d9d3bd7b0bdafdbfadee8 plugrl_env_client/envs/classic/classic_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/classic/__init__.py +616812fa3e47727ecf32d58193ea82d7491fa72e7248c14693657120c61f53e4 plugrl_env_client/envs/d4rl/d4rl_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/d4rl/__init__.py +e0461287f53363eabbb8503418e0de20c93279cdb865e6eb8b4016ba56bd51fd plugrl_env_client/envs/dummy_env.py +2162e3cf76b75ffe347adbf39dbd87e23eab36117cce8626d6145049683cbb9e plugrl_env_client/envs/__init__.py +4f308a45cc7c78c3617d172ccc60305efebba82605b6f32737ac75c8a6ea5ddc plugrl_env_client/envs/libero/image_tools.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/libero/__init__.py +f835ff5b01161df7c8c83ad76f06a7851c5306c2ccb1b35721db0a470552e0f8 plugrl_env_client/envs/libero/libero_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/mujoco/__init__.py +ea270ca84b88c4857ad02bcdf3dce7320057b85962f149d06a9261b6cfa0c862 plugrl_env_client/envs/mujoco/mujoco_env.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/envs/robomimic/__init__.py +22c384593ac46df55c83812bfe9697c86c75e9a5cf487c09458bd5fbe120a353 plugrl_env_client/envs/robomimic/robomimic_env.py +8435d17ff7f893138a198c0452abbd1d98a51bb80a0cac362801964c280f6844 plugrl_env_client/__init__.py +1c07862ce6a46e8c25e271d0061f0b82333504d171e78a280188ea394f9308cb plugrl_env_client/recorder/args.py +d3c58231640d9bce38619c5c6d56e1d348177b6dbce5537f8a4fd10c382df883 plugrl_env_client/recorder/events.py +95b5fa86e4b13985b58e79bc1605fc91aed42418db67ba0e64b85cde0f93f943 plugrl_env_client/recorder/__init__.py +2f2a6a94ae4566a95bfece985572eddb51aeef8335bd8fe67fc5753417deae5c plugrl_env_client/recorder/recorder.py +b8d0b07d090eb0f97a50e7f7e542376b63cbd507dcd8ffb79fcaec1f1c744dc5 plugrl_env_client/recorder/sink.py +e3e42f2354b9f83eb15f99ad537f389cc29b4411fd173cf676c711037dea1176 plugrl_env_client/recorder/writers.py +e5d0ed0b9a141e762ff85dc96471a495c85bf9869084192518221c4c86ae61c6 plugrl_env_client/runner/args.py +fb6a4761fb3bd68440546ba1f4ace66ad4be1014ac0b5668041c06e3d2eac117 plugrl_env_client/runner/__init__.py +257b347e940daeefb025f3fa3ecca7316849a50d537bdd313473c86d692e998c plugrl_env_client/runner/rollout.py +66cbdf5d1572c987bfedd3ffad896aafb4b8f865dcbc6876e7e8037193367607 plugrl_env_client/runner/run.py +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 plugrl_env_client/utils/__init__.py +76908794d381ef84755471be61ecc86bfe973eb42501541ce7815d786156df7f plugrl_env_client/utils/recorder.py +e75f1539456ac6b2054d0a3d7fa5886f672783b75975da111b86daacc5cb8bb6 plugrl_env_client/utils/registration.py +99c3587c966a918b83a65de63556fb58e24761055e07ecf32a7f269be268ec5a plugrl_env_client/utils/rollout.py +1a5766bb33b2e43472980c9ef7e76367c7ff06aeeafd4db1fe5cf134894772e7 plugrl_env_client/utils/wrappers/episode_stats_wrapper.py +65ebc060d3c5c1898ef54298e5e0e9588431ada249a90583180fbe1056fddb4e plugrl_env_client/utils/wrappers/__init__.py +e246f2ac99b771225302b87952a4050a7514e9baeb2dba5a1240af50f98ec240 plugrl_env_client/utils/wrappers/real_time_wrapper.py +827c8f87b7f56396885a1222d02d098185af675278b452dcc2124c49cbe1b9e9 plugrl_env_client/utils/wrappers/time_limit_wrapper.py diff --git a/experiments/e36-pi0-longer/results/learn-stats.txt b/experiments/e36-pi0-longer/results/learn-stats.txt new file mode 100644 index 0000000..7d54e63 --- /dev/null +++ b/experiments/e36-pi0-longer/results/learn-stats.txt @@ -0,0 +1,26 @@ +## fpopp-s7 +/home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/tensorboard/events.out.tfevents.1790515737.e88d1a355314.38426.0 + fpo/policy_ratio_mean 1 0.9251 0.9428 0.9021 0.923 0.8417 0.8406 0.8189 0.4331 + fpo/clipped_ratio_mean 0 0.4562 0.5206 0.4626 0.4264 0.5009 0.579 0.754 0.9633 + fpo/initial_cfm_loss_mean 0.1147 0.1306 0.1622 0.2615 0.8139 5.848 28.08 73.75 114.4 + fpo/cfm_loss_mean 0.1147 0.2441 0.2526 0.954 6.069 30.24 74.4 151.5 239.8 + rollout/success 0.6 0.381 0.7333 0.4048 0.09756 0.1429 0.05 0 0 + rollout/reward 0.6 0.381 0.7333 0.4048 0.09756 0.1429 0.05 0 0 + rollout/length 458.1 483.3 438.7 482.6 509.6 504.2 512.4 520 520 +## fpopp-s8 +/home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/tensorboard/events.out.tfevents.1790515793.e88d1a355314.39189.0 + fpo/policy_ratio_mean 1 0.9292 0.8118 0.8335 0.7735 0.798 0.6176 0.3832 0.8228 + fpo/clipped_ratio_mean 0 0.4927 0.4774 0.4361 0.5372 0.534 0.6457 0.992 0.6372 + fpo/initial_cfm_loss_mean 0.1196 0.1226 0.1855 1.895 9.969 51.65 179.5 338.3 406.9 + fpo/cfm_loss_mean 0.1196 0.2307 5.222 12.02 62.97 191.6 362.6 527.9 924.2 + rollout/success 0.6667 0.5 0.4651 0.175 0.125 0 0 0 0 + rollout/reward 0.6667 0.5 0.4651 0.175 0.125 0 0 0 0 + rollout/length 461.2 470 475.2 510.2 512.6 520 520 520 520 +## dppo +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/tensorboard/events.out.tfevents.1790515856.e88d1a355314.40900.0 + losses/old_approx_kl 5.379e-05 -0.0001036 8.709e-05 0.0002731 0.0001768 0.0002059 -3.392e-05 7.548e-05 0.0001538 0.0001972 + losses/approx_kl 2.563e-07 2.224e-07 2.941e-07 2.926e-07 1.284e-07 3.461e-07 2.403e-07 1.982e-07 1.721e-07 4.135e-07 + losses/clipfrac 0.07959 0.08201 0.08027 0.08383 0.08203 0.09467 0.08831 0.092 0.08889 0.092 + rollout/success 0.6 0.5349 0.6591 0.5682 0.6364 0.7551 0.561 0.6087 0.6279 0.4762 + rollout/reward 0.6 0.5349 0.6591 0.5682 0.6364 0.7551 0.561 0.6087 0.6279 0.4762 + rollout/length 467.7 465.2 450.3 474.9 457.6 442 465.9 465.6 460.1 478.7 diff --git a/experiments/e36-pi0-longer/results/movement.txt b/experiments/e36-pi0-longer/results/movement.txt new file mode 100644 index 0000000..c146324 --- /dev/null +++ b/experiments/e36-pi0-longer/results/movement.txt @@ -0,0 +1,24 @@ +/home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/20489/model.safetensors + mod moved 74 unchanged 0 relative distance 0.041006 + attn moved 72 unchanged 0 relative distance 0.030011 + mlp moved 54 unchanged 0 relative distance 0.029143 + io moved 8 unchanged 0 relative distance 0.023628 + expert moved 200 unchanged 1 relative distance 0.028341 +/home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/20480/model.safetensors + mod moved 74 unchanged 0 relative distance 0.040266 + attn moved 72 unchanged 0 relative distance 0.027416 + mlp moved 54 unchanged 0 relative distance 0.028762 + io moved 8 unchanged 0 relative distance 0.027476 + expert moved 200 unchanged 1 relative distance 0.027503 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/20490/model.safetensors + mod moved 74 unchanged 0 relative distance 0.004403 + attn moved 72 unchanged 0 relative distance 0.000465 + mlp moved 54 unchanged 0 relative distance 0.000472 + io moved 8 unchanged 0 relative distance 0.004367 + expert moved 200 unchanged 1 relative distance 0.001622 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/40970/model.safetensors + mod moved 74 unchanged 0 relative distance 0.006455 + attn moved 72 unchanged 0 relative distance 0.000613 + mlp moved 54 unchanged 0 relative distance 0.000629 + io moved 8 unchanged 0 relative distance 0.006365 + expert moved 200 unchanged 1 relative distance 0.002366 diff --git a/experiments/e36-pi0-longer/results/run.log b/experiments/e36-pi0-longer/results/run.log new file mode 100644 index 0000000..b6737a6 --- /dev/null +++ b/experiments/e36-pi0-longer/results/run.log @@ -0,0 +1,48 @@ +[09-27 13:28:41] code: 5272832 (feat/fpo-chunk-loss-per-sample-ratio, PR #74: main 2bd2789 plus FPO++ chunk loss and per-sample ratio), LF archive +[09-27 13:28:41] --- phase 1 +[09-28 00:08:55] END train dppo rc=0 +[09-28 01:31:18] END train fpopp-s7 rc=0 +[09-28 01:32:32] END train fpopp-s8 rc=0 +[09-28 01:32:32] checkpoints fpopp-s7: 4100 8200 12290 16390 20489 24579 28679 32769 36869 +[09-28 01:32:32] checkpoints fpopp-s8: 4100 8200 12290 16390 20480 24580 28679 32769 36869 +[09-28 01:32:32] checkpoints dppo: 4100 8200 12290 16390 20490 24580 28680 32770 36870 40970 +[09-28 01:32:32] --- phase 2 +[09-28 01:32:32] --- movement from the base, iterations 5 and 10 +[09-28 01:32:32] END eval fpopp-s8-it10 not run: no checkpoint +[09-28 01:32:32] END eval fpopp-s7-it10 not run: no checkpoint +/home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/20489/model.safetensors + mod moved 74 unchanged 0 relative distance 0.041006 + attn moved 72 unchanged 0 relative distance 0.030011 + mlp moved 54 unchanged 0 relative distance 0.029143 + io moved 8 unchanged 0 relative distance 0.023628 + expert moved 200 unchanged 1 relative distance 0.028341 +/home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/20480/model.safetensors + mod moved 74 unchanged 0 relative distance 0.040266 + attn moved 72 unchanged 0 relative distance 0.027416 + mlp moved 54 unchanged 0 relative distance 0.028762 + io moved 8 unchanged 0 relative distance 0.027476 + expert moved 200 unchanged 1 relative distance 0.027503 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/20490/model.safetensors + mod moved 74 unchanged 0 relative distance 0.004403 + attn moved 72 unchanged 0 relative distance 0.000465 + mlp moved 54 unchanged 0 relative distance 0.000472 + io moved 8 unchanged 0 relative distance 0.004367 + expert moved 200 unchanged 1 relative distance 0.001622 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/40970/model.safetensors + mod moved 74 unchanged 0 relative distance 0.006455 + attn moved 72 unchanged 0 relative distance 0.000613 + mlp moved 54 unchanged 0 relative distance 0.000629 + io moved 8 unchanged 0 relative distance 0.006365 + expert moved 200 unchanged 1 relative distance 0.002366 +[09-28 01:34:56] --- learn-step statistics +[09-28 02:16:17] END eval dppo-it5 rc=0 +[09-28 02:16:47] END eval dppo-it10 rc=0 +[09-28 02:18:27] END eval fpopp-s7-it5 rc=0 +[09-28 02:18:53] END eval fpopp-s8-it5 rc=0 +[09-28 02:18:53] === rows === +policy task_id episodes successes sr ci_lo ci_hi valid +dppo-it5 8 50 28 0.5600 0.4231 0.6884 true +dppo-it10 8 50 20 0.4000 0.2761 0.5382 true +fpopp-s7-it5 8 50 5 0.1000 0.0435 0.2136 true +fpopp-s8-it5 8 50 0 0.0000 0.0000 0.0713 true +[09-28 02:18:53] E36_DONE diff --git a/experiments/e36-pi0-longer/results/run.out b/experiments/e36-pi0-longer/results/run.out new file mode 100644 index 0000000..b6737a6 --- /dev/null +++ b/experiments/e36-pi0-longer/results/run.out @@ -0,0 +1,48 @@ +[09-27 13:28:41] code: 5272832 (feat/fpo-chunk-loss-per-sample-ratio, PR #74: main 2bd2789 plus FPO++ chunk loss and per-sample ratio), LF archive +[09-27 13:28:41] --- phase 1 +[09-28 00:08:55] END train dppo rc=0 +[09-28 01:31:18] END train fpopp-s7 rc=0 +[09-28 01:32:32] END train fpopp-s8 rc=0 +[09-28 01:32:32] checkpoints fpopp-s7: 4100 8200 12290 16390 20489 24579 28679 32769 36869 +[09-28 01:32:32] checkpoints fpopp-s8: 4100 8200 12290 16390 20480 24580 28679 32769 36869 +[09-28 01:32:32] checkpoints dppo: 4100 8200 12290 16390 20490 24580 28680 32770 36870 40970 +[09-28 01:32:32] --- phase 2 +[09-28 01:32:32] --- movement from the base, iterations 5 and 10 +[09-28 01:32:32] END eval fpopp-s8-it10 not run: no checkpoint +[09-28 01:32:32] END eval fpopp-s7-it10 not run: no checkpoint +/home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/20489/model.safetensors + mod moved 74 unchanged 0 relative distance 0.041006 + attn moved 72 unchanged 0 relative distance 0.030011 + mlp moved 54 unchanged 0 relative distance 0.029143 + io moved 8 unchanged 0 relative distance 0.023628 + expert moved 200 unchanged 1 relative distance 0.028341 +/home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/20480/model.safetensors + mod moved 74 unchanged 0 relative distance 0.040266 + attn moved 72 unchanged 0 relative distance 0.027416 + mlp moved 54 unchanged 0 relative distance 0.028762 + io moved 8 unchanged 0 relative distance 0.027476 + expert moved 200 unchanged 1 relative distance 0.027503 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/20490/model.safetensors + mod moved 74 unchanged 0 relative distance 0.004403 + attn moved 72 unchanged 0 relative distance 0.000465 + mlp moved 54 unchanged 0 relative distance 0.000472 + io moved 8 unchanged 0 relative distance 0.004367 + expert moved 200 unchanged 1 relative distance 0.001622 +/home/gotham/tmp/plugrl/e36/runs/dppo/ck/dppo/pi0-policy/dppo/40970/model.safetensors + mod moved 74 unchanged 0 relative distance 0.006455 + attn moved 72 unchanged 0 relative distance 0.000613 + mlp moved 54 unchanged 0 relative distance 0.000629 + io moved 8 unchanged 0 relative distance 0.006365 + expert moved 200 unchanged 1 relative distance 0.002366 +[09-28 01:34:56] --- learn-step statistics +[09-28 02:16:17] END eval dppo-it5 rc=0 +[09-28 02:16:47] END eval dppo-it10 rc=0 +[09-28 02:18:27] END eval fpopp-s7-it5 rc=0 +[09-28 02:18:53] END eval fpopp-s8-it5 rc=0 +[09-28 02:18:53] === rows === +policy task_id episodes successes sr ci_lo ci_hi valid +dppo-it5 8 50 28 0.5600 0.4231 0.6884 true +dppo-it10 8 50 20 0.4000 0.2761 0.5382 true +fpopp-s7-it5 8 50 5 0.1000 0.0435 0.2136 true +fpopp-s8-it5 8 50 0 0.0000 0.0000 0.0713 true +[09-28 02:18:53] E36_DONE diff --git a/experiments/e36-pi0-longer/results/stageC_eval.tsv b/experiments/e36-pi0-longer/results/stageC_eval.tsv new file mode 100644 index 0000000..5e3f8f2 --- /dev/null +++ b/experiments/e36-pi0-longer/results/stageC_eval.tsv @@ -0,0 +1,5 @@ +policy task_id episodes successes sr ci_lo ci_hi valid +dppo-it5 8 50 28 0.5600 0.4231 0.6884 true +dppo-it10 8 50 20 0.4000 0.2761 0.5382 true +fpopp-s7-it5 8 50 5 0.1000 0.0435 0.2136 true +fpopp-s8-it5 8 50 0 0.0000 0.0000 0.0713 true diff --git a/experiments/e36-pi0-longer/results/stageC_eval_correctness.tsv b/experiments/e36-pi0-longer/results/stageC_eval_correctness.tsv new file mode 100644 index 0000000..8c78cc6 --- /dev/null +++ b/experiments/e36-pi0-longer/results/stageC_eval_correctness.tsv @@ -0,0 +1,5 @@ +cell policy client_exit wall_s client_episodes client_steps server_episodes server_length_sum episode_accounting step_accounting no_step_state_warnings reconnects valid +e36-dppo-it5 dppo-it5 0 1998 50 23398 50 23398 true true 0 0 true +e36-dppo-it10 dppo-it10 0 2024 50 23927 50 23927 true true 0 0 true +e36-fpopp-s7-it5 fpopp-s7-it5 0 2126 50 25234 50 25234 true true 0 0 true +e36-fpopp-s8-it5 fpopp-s8-it5 0 2148 50 26000 50 26000 true true 0 0 true diff --git a/experiments/e36-pi0-longer/results/train-dppo.out b/experiments/e36-pi0-longer/results/train-dppo.out new file mode 100644 index 0000000..8d87821 --- /dev/null +++ b/experiments/e36-pi0-longer/results/train-dppo.out @@ -0,0 +1,8 @@ +[09-27 13:30:41] cell=dppo iters=10 buffer=4096 batch=8 port=8473 src=5272832 (feat/fpo-chunk-loss-per-sample-ratio, PR #74: main 2bd2789 plus FPO++ chunk loss and per-sample ratio), LF archive +[09-27 13:30:41] ALLOW_SIBLINGS=1: not checking for other clients or servers +[09-27 13:33:16] server listening +[09-28 00:08:53] server exited after 38137s +[09-28 00:08:55] after teardown: 28 client or server processes on the machine +[09-28 00:08:55] checkpoints: dppo/pi0-policy/dppo/12290/model.safetensors dppo/pi0-policy/dppo/16390/model.safetensors dppo/pi0-policy/dppo/20490/model.safetensors dppo/pi0-policy/dppo/24580/model.safetensors dppo/pi0-policy/dppo/28680/model.safetensors dppo/pi0-policy/dppo/32770/model.safetensors dppo/pi0-policy/dppo/36870/model.safetensors dppo/pi0-policy/dppo/40970/model.safetensors dppo/pi0-policy/dppo/4100/model.safetensors dppo/pi0-policy/dppo/8200/model.safetensors +[09-28 00:08:55] peak MiB on the server cards 4 and 5: 18520, 2802 +[09-28 00:08:55] E36_DPPO_DONE dppo diff --git a/experiments/e36-pi0-longer/results/train-fpopp-s7.out b/experiments/e36-pi0-longer/results/train-fpopp-s7.out new file mode 100644 index 0000000..bc13584 --- /dev/null +++ b/experiments/e36-pi0-longer/results/train-fpopp-s7.out @@ -0,0 +1,26 @@ +[13:28:41] cell=fpopp-s7 iterations=10 master_device=cuda:1 lr=1e-5 updates=4 clip=0.05 drift=0 warmup=0 srv_gpus=0,1 cli_gpu=3 buffer=4096 batch=8 +[13:28:41] ALLOW_SIBLINGS=1: not checking for other clients or servers +[13:28:41] arm flags: --algo.n-critic-warmup-itrs 1 --algo.output-mode u --algo.no-discretize-t-for-training --algo.cfm-loss-steps 5 --algo.cfm-loss-dims 7 --algo.cfm-loss-sum-over-steps --algo.ratio-per-sample +[13:30:36] server listening +[01:30:44] clients exit 124 after 43208s (server_died=0) +/home/gotham/tmp/plugrl/e36/train.sh: line 80: 38426 Killed ( cd "$R/plugrl-server-e32" && exec setsid env PYTHONPATH="$R/plugrl-server-e32/src" OPENPI_DATA_HOME="$R/.cache/openpi" XDG_CACHE_HOME="$R/.cache" TMPDIR="$R/.tmp" HF_HOME="$R/.cache/hf" TORCHINDUCTOR_CACHE_DIR="$R/.cache/inductor" CUDA_VISIBLE_DEVICES=${SRV_GPUS:-0,1} PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True "$SPY" -m plugrl_server.cli pi0-policy default fpo default --policy.name pi05_libero --policy.checkpoint-path "$R/ckpt/pi05_libero" --policy.device cuda:0 --algo.learning-rate ${LR:-1e-5} --algo.batch-size "$BATCH_SIZE" --algo.num-updates-per-batch ${UPDATES:-4} --algo.n-samples-per-action "$N_SAMPLES" --algo.buffer-size "$BUFFER_SIZE" --algo.clipping-epsilon ${CLIP_EPS:-0.05} --algo.global-steps "$GLOBAL_STEPS" --algo.save-interval 1 "${MASTER_FLAG[@]}" "${OPT_FLAGS[@]}" --seed ${SEED:-7} --port "$PORT" --no-show-progress-bar --no-show-metric-table --checkpoint-base-dir "$OUT/ck" --exp-name "$CELL" --overwrite ) > "$OUT/server.log" 2>&1 +[01:31:18] after teardown, clients: 12 alive + PID ELAPSED COMMAND +43113 43164 /home/gotham/tmp/plugrl/venv-libero/bin/python -m plugrl_env_client.cli libero-v1 --server-host 127.0.0.1 +43212 43160 /home/gotham/tmp/plugrl/venv-libero/bin/python -c from multiprocessing.resource_tracker import main;main(4 +43213 43160 /home/gotham/tmp/plugrl/venv-libero/bin/python -c from multiprocessing.spawn import spawn_main; spawn_main +43214 43160 /home/gotham/tmp/plugrl/venv-libero/bin/python -c from multiprocessing.spawn import spawn_main; spawn_main +[01:31:18] TEARDOWN_NOT_CLEAN +[01:31:18] === learn steps seen === +48 +[01:31:18] === out of memory? === +0 +[01:31:18] === peak memory, GPU0 and GPU1 === +gpu0_peak_mib=20082 gpu1_peak_mib=9070 +[01:31:18] === checkpoints === +18:36 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/16390/model.safetensors +19:52 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/20489/model.safetensors +21:07 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/24579/model.safetensors +22:23 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/28679/model.safetensors +23:39 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s7/ck/fpo/pi0-policy/fpopp-s7/32769/model.safetensors +[01:31:18] E36_TRAIN_DONE fpopp-s7 diff --git a/experiments/e36-pi0-longer/results/train-fpopp-s8.out b/experiments/e36-pi0-longer/results/train-fpopp-s8.out new file mode 100644 index 0000000..efacb24 --- /dev/null +++ b/experiments/e36-pi0-longer/results/train-fpopp-s8.out @@ -0,0 +1,19 @@ +[13:29:41] cell=fpopp-s8 iterations=10 master_device=cuda:1 lr=1e-5 updates=4 clip=0.05 drift=0 warmup=0 srv_gpus=2,3 cli_gpu=0 buffer=4096 batch=8 +[13:29:41] ALLOW_SIBLINGS=1: not checking for other clients or servers +[13:29:41] arm flags: --algo.n-critic-warmup-itrs 1 --algo.output-mode u --algo.no-discretize-t-for-training --algo.cfm-loss-steps 5 --algo.cfm-loss-dims 7 --algo.cfm-loss-sum-over-steps --algo.ratio-per-sample +[13:31:53] server listening +[01:31:59] clients exit 124 after 43206s (server_died=0) +[01:32:32] teardown clean +[01:32:32] === learn steps seen === +44 +[01:32:32] === out of memory? === +0 +[01:32:32] === peak memory, GPU0 and GPU1 === +gpu0_peak_mib=20082 gpu1_peak_mib=9070 +[01:32:32] === checkpoints === +18:31 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/16390/model.safetensors +19:45 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/20480/model.safetensors +21:00 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/24580/model.safetensors +22:14 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/28679/model.safetensors +23:28 8528818570 /home/gotham/tmp/plugrl/e36/fpopp-s8/ck/fpo/pi0-policy/fpopp-s8/32769/model.safetensors +[01:32:32] E36_TRAIN_DONE fpopp-s8