Skip to content

docs: algorithm, policy and contributing pages match the code - #24

Merged
tactino merged 1 commit into
mainfrom
docs/algorithm-policy
Sep 29, 2026
Merged

tactino merged 1 commit into
mainfrom
docs/algorithm-policy

Conversation

@tactino

@tactino tactino commented Sep 29, 2026

Copy link
Copy Markdown
Member

An audit of the algorithm, policy and contributing pages against plugrl-server found 21 discrepancies. This fixes them in both languages.

Wrong

  • plugrl-run-server --help lists policies only. Algorithms appear under <policy> <variant> --help.
  • Six algorithms are registered, not five. ppo was missing entirely, and now has a subsection: CleanRL's PPO, the default and dppo-square variants, and which policies it takes.
  • No algorithm needs the dppo extra. dppo-policy and dppo-gaussian-policy do. Without it they are silently absent from the CLI; the quoted log line did not exist.
  • pi0-policy does run under dppo, with the libero variant, as E25 did.
  • --policy.train-expert-only true is a parse error. The pages now use the flag form.
  • The E11 command now matches E11's script.

Stale. The training loop page now describes the server after #104 and #106:

  • no inference while the algorithm wants to learn;
  • frames from an older policy, or arriving while it waits to learn, go to discard_feedback and are counted in server/discarded_frames;
  • global_step counts stored frames;
  • the rule applies to on-policy algorithms, with on_policy = False for replay buffers.

The page also now covers resuming: the directory layout, FileNotFoundError and FileExistsError, and the Ctrl-C checkpoint. It covers warm starts from another run's weights (--algo.policy-checkpoint-path, --algo.restore) as well.

Missing

  • Custom algorithm contract: discard_feedback (an override must call super()), on_policy, derive_train_state, pre_learn/post_learn, init_optimizers, and load_checkpoint on --resume. Registration needs both the config module and the class module imported.
    • The template now uses global_steps, and calls record_episode_metrics, without which rollout/* stays 0.
  • Policies: gaussian-policy and dppo-gaussian-policy; batch-first (B, H, ...) actions; action_dim/action_horizon for the metadata; a runtime state that can be sliced per env; and the critic DPPO and FPO require.
  • Flow policies: FPO's base, _predict_v and dt.
  • InternalState no longer exists; the pages name the real types.

Checked. Every plugrl-run-server command on these pages parses against the CLI. pi0-policy was checked with its OpenPI imports stubbed out, as OpenPI is not installed here. The custom algorithm template was loaded and each method called on dummy-policy. mkdocs build --strict passes.

Merge after plugrl-server #106 (on_policy) and #107 (the dummy policy's defaults), which these pages describe.

…ix algorithms, the frame rule, resuming, and the full algorithm and policy contracts
@tactino
tactino merged commit 4f52ad3 into main Sep 29, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant