feat(dppo): make the log-probability clamp configurable - #64
Merged
Merged
Conversation
DPPO clamps every element's log-probability to [-5, 2] before forming the ratio, as its reference does. With dppo-policy the upper bound never binds (every step's deviation is at least sampling_noise_level; at 0.1 the log-density is at most 1.38). fpo-policy's flow steps are narrower - 0.1 down to about 0.01 at the same level - and 50.7% of its elements land above 2 (up to 93% at the last step), where they carry no gradient and a ratio of exactly 1. The bounds are now logprob_clamp_min/max, defaulting to the reference's, so the change is inert until someone lifts them. Tests written first (5 of 5 failing before): the defaults, the fraction on a fresh fpo-policy at E28's level, a fully clamped batch has zero actor gradient, lifting the bounds restores it, and the flags parse (-inf needs the equals form).
tactino
added a commit
that referenced
this pull request
Sep 26, 2026
…op fpo-policy under DPPO Three arms of E28's fpodppo-hopper: control, noclamp (upper bound lifted, #64), wide (noise level 1.0). V1: logged entropy -1.920 / +0.383. P2 the control stays within +50; P3 noclamp and P4 wide gain >= +100 on 2 of 3, both expected falsified after the pilot. Pilot, the clamp probe and the actor movement of E28's cells recorded before any registered run.
Member
Author
|
E29 (#65) tested this: lifting the upper bound on E28's |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DPPO's loss clamps every element's log-probability to [-5, 2] before forming the ratio - DPPO's reference values. This makes the bounds configuration, defaulting to exactly those values, so nothing changes until someone lifts them.
Why it matters
With
dppo-policythe upper bound never binds: each denoising step's deviation is floored atsampling_noise_level, and a normal with deviation 0.1 has a log-density of at most 1.38.fpo-policyis a flow policy; its step deviation is σ_t·√dt with σ_t = level·√(t/(1−t)), which at level 0.1 runs from 0.1 at the first of its ten steps down to about 0.01 at the last. Measured on freshly initialised policies at E28's settings (Hopper, level 0.1, 4,096 samples):fpo-policydppo-policyThe entropies are exactly what E28's tensorboards logged for these cells. A clamped element carries no gradient and contributes a ratio of exactly 1, which fits what E18 and E28 (#63) logged for
fpo-policyunder DPPO:approx_klaround 10⁻⁷ againstdppo-policy's 10⁻⁵,clipfracnear 0, and no learning on HalfCheetah, Hopper or Walker2d - while the same policy learns all three under FPO. Whether the clamp is why is not settled here; E29 will test it against the narrow noise itself.Change
DPPOAlgoConfig.logprob_clamp_min = -5.0,logprob_clamp_max = 2.0;_compute_lossreads them for both the new and old log-probabilities.DPPOAlgoDistributedConfiginherits them.inflifts a bound.--algo.logprob-clamp-min=-infneeds the equals form: argparse reads a bare-infas a flag (the test pins this).Tests
tests/test_dppo_logprob_clamp.py, written first (5 of 5 failing before):fpo-policyat level 0.1, 35-65% of elements are above the upper bound;Full suite locally: 186 passed, 3 skipped, 6 xfailed, 4 xpassed.