exp: E29 - the logprob clamp is not why fpo-policy stands still under DPPO; its narrow noise is part of it - #65
Merged
Merged
Conversation
DPPO clamps every element's log-probability to [-5, 2] before forming the ratio, as its reference does. With dppo-policy the upper bound never binds (every step's deviation is at least sampling_noise_level; at 0.1 the log-density is at most 1.38). fpo-policy's flow steps are narrower - 0.1 down to about 0.01 at the same level - and 50.7% of its elements land above 2 (up to 93% at the last step), where they carry no gradient and a ratio of exactly 1. The bounds are now logprob_clamp_min/max, defaulting to the reference's, so the change is inert until someone lifts them. Tests written first (5 of 5 failing before): the defaults, the fraction on a fresh fpo-policy at E28's level, a fully clamped batch has zero actor gradient, lifting the bounds restores it, and the flags parse (-inf needs the equals form).
E24's run_cell.sh plus EXTRA_ALGO_ARGS; run.sh runs E28's fpodppo-hopper cell as three arms that differ only in those flags: control, noclamp (--algo.logprob-clamp-max inf, #64) and wide (noise level 1.0 for sampling and log-probability).
…op fpo-policy under DPPO Three arms of E28's fpodppo-hopper: control, noclamp (upper bound lifted, #64), wide (noise level 1.0). V1: logged entropy -1.920 / +0.383. P2 the control stays within +50; P3 noclamp and P4 wide gain >= +100 on 2 of 3, both expected falsified after the pilot. Pilot, the clamp probe and the actor movement of E28's cells recorded before any registered run.
… but misses the bar P1 3/3, V1 pass, P2 holds (control +10.6/+4.5/+9.7, E28's to a point). P3 falsified: lifting the clamp gives +7.0/+3.0/+7.8. P4 falsified by its rule: ten times the noise gives +73.6/+84.6/+106.2, one seed past +100, seven to nineteen times the control and still climbing. The pilot's approx_kl was misread: across noise levels it scales as 1/sigma^2.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fpo-policylearns Hopper under FPO and not under DPPO (E28, #63), wheredppo-policygains +260 to +417. Two code-level differences where the policies meet DPPO's loss: the log-probability clamp [-5, 2] binds on 50.7% offpo-policy's elements and none ofdppo-policy's, and its per-step noise is about a tenth. E29 changes each on E28's cell. Stacked on #64, which makes the clamp configurable (thenoclamparm needs it).Result (gain = mean of iterations 91-100 minus the first)
controlnoclamp--algo.logprob-clamp-max infwidewidegains 7-19× its control seed, flat for ~40 iterations then climbing, still climbing at 100 (episodes 23 → 58-76 steps). The registered bar was not met, so "neither is sufficient" is what is claimed. At matched entropy it reaches about a quarter ofdppo-policy's gain; network, chunk (1 vs 4), denoising steps (10 vs 20) and the output scale remain untested.approx_kl: across noise levels it scales as 1/σ² for the same shift, andwide's actor actually moved 40-43% against the control's 20-23%.Files
PROTOCOL.md(pre-registered,20484bf, after the pilot),FINDINGS.md,results/pilot.txtrun.sh,run_cell.sh(E24's plusEXTRA_ALGO_ARGS),summarise.py,clamp_probe.pysummary.tsv,results/- every server and client log,verdicts.txt; checkpoints stay on the workstation