Skip to content

feat(dppo): make the log-probability clamp configurable - #64

Merged
tactino merged 1 commit into
mainfrom
feat/dppo-logprob-clamp
Sep 27, 2026
Merged

tactino merged 1 commit into
mainfrom
feat/dppo-logprob-clamp

Conversation

@tactino

@tactino tactino commented Sep 26, 2026

Copy link
Copy Markdown
Member

DPPO's loss clamps every element's log-probability to [-5, 2] before forming the ratio - DPPO's reference values. This makes the bounds configuration, defaulting to exactly those values, so nothing changes until someone lifts them.

Why it matters

With dppo-policy the upper bound never binds: each denoising step's deviation is floored at sampling_noise_level, and a normal with deviation 0.1 has a log-density of at most 1.38. fpo-policy is a flow policy; its step deviation is σ_t·√dt with σ_t = level·√(t/(1−t)), which at level 0.1 runs from 0.1 at the first of its ten steps down to about 0.01 at the last. Measured on freshly initialised policies at E28's settings (Hopper, level 0.1, 4,096 samples):

log-prob elements above 2 mean entropy
fpo-policy 50.7% - 0 at t = 1.0-0.8, then 0.37, 0.59, 0.70, 0.77, 0.83, 0.88, 0.93 at t = 0.1 −1.920
dppo-policy 0% (0.17% below −5) 0.312

The entropies are exactly what E28's tensorboards logged for these cells. A clamped element carries no gradient and contributes a ratio of exactly 1, which fits what E18 and E28 (#63) logged for fpo-policy under DPPO: approx_kl around 10⁻⁷ against dppo-policy's 10⁻⁵, clipfrac near 0, and no learning on HalfCheetah, Hopper or Walker2d - while the same policy learns all three under FPO. Whether the clamp is why is not settled here; E29 will test it against the narrow noise itself.

Change

  • DPPOAlgoConfig.logprob_clamp_min = -5.0, logprob_clamp_max = 2.0; _compute_loss reads them for both the new and old log-probabilities. DPPOAlgoDistributedConfig inherits them.
  • inf lifts a bound. --algo.logprob-clamp-min=-inf needs the equals form: argparse reads a bare -inf as a flag (the test pins this).

Tests

tests/test_dppo_logprob_clamp.py, written first (5 of 5 failing before):

  • the defaults are the reference's bounds;
  • on a fresh fpo-policy at level 0.1, 35-65% of elements are above the upper bound;
  • a batch clamped everywhere gives the actor exactly zero policy gradient;
  • lifting the bounds restores a larger, non-zero gradient;
  • the bounds parse from the command line.

Full suite locally: 186 passed, 3 skipped, 6 xfailed, 4 xpassed.

DPPO clamps every element's log-probability to [-5, 2] before forming the
ratio, as its reference does. With dppo-policy the upper bound never binds
(every step's deviation is at least sampling_noise_level; at 0.1 the
log-density is at most 1.38). fpo-policy's flow steps are narrower - 0.1
down to about 0.01 at the same level - and 50.7% of its elements land above
2 (up to 93% at the last step), where they carry no gradient and a ratio of
exactly 1. The bounds are now logprob_clamp_min/max, defaulting to the
reference's, so the change is inert until someone lifts them.

Tests written first (5 of 5 failing before): the defaults, the fraction on
a fresh fpo-policy at E28's level, a fully clamped batch has zero actor
gradient, lifting the bounds restores it, and the flags parse (-inf needs
the equals form).
tactino added a commit that referenced this pull request Sep 26, 2026
…op fpo-policy under DPPO

Three arms of E28's fpodppo-hopper: control, noclamp (upper bound lifted,
#64), wide (noise level 1.0). V1: logged entropy -1.920 / +0.383. P2 the
control stays within +50; P3 noclamp and P4 wide gain >= +100 on 2 of 3,
both expected falsified after the pilot. Pilot, the clamp probe and the
actor movement of E28's cells recorded before any registered run.
@tactino

tactino commented Sep 27, 2026

Copy link
Copy Markdown
Member Author

E29 (#65) tested this: lifting the upper bound on E28's fpo-policy · DPPO · Hopper cell gives +7.0 / +3.0 / +7.8 against the control's +10.6 / +4.5 / +9.7 - no effect. The clamp binds as described, but it is not why fpo-policy stands still under DPPO; ten times the noise moves it (+74 to +106). So this PR stays what it is - the bounds configurable, the reference's defaults unchanged - and E29 gives no reason to change the default. E29 needs it merged to be reproducible.

@tactino
tactino merged commit 865dbf7 into main Sep 27, 2026
3 checks passed
@tactino
tactino deleted the feat/dppo-logprob-clamp branch September 27, 2026 06:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant