Skip to content

E41: E39 resumed to FPO++'s 8M steps - fpo-policy · FPO learns square by the bar, two of three seeds - #96

Merged
tactino merged 4 commits into
mainfrom
exp/e41-square-fpo-plus-plus-8m
Sep 28, 2026
Merged

tactino merged 4 commits into
mainfrom
exp/e41-square-fpo-plus-plus-8m

Conversation

@tactino

@tactino tactino commented Sep 28, 2026

Copy link
Copy Markdown
Member

E39's fpo-policy · FPO · square cell, resumed from its final checkpoints and run on to FPO++'s own 8M steps (167 iterations), as pre-registered in PROTOCOL.md before the run.

By the registered rule the cell learns: over iterations 158-167, seeds 0 and 1 are +0.291 and +0.214 above their start, and seed 2 is +0.153. P1 and V1 hold, and P2 holds. Its grounds were wrong about which seed would cross: seed 1 did and seed 2 fell back.

What the findings flag: every seed came down from its best window over the last thirty to sixty iterations (seed 2 by 0.09). Fifty-episode evaluations of the final checkpoints are 0.82 / 0.64 / 0.54, against the clone's 0.50 and E39's 0.80 / 0.64 / 0.64.

Only files under experiments/e41-square-fpo-plus-plus-8m/.

@tactino
tactino merged commit 714ae6a into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant