docs: algorithm, policy and contributing pages match the code - #24
Merged
Merged
Conversation
…ix algorithms, the frame rule, resuming, and the full algorithm and policy contracts
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An audit of the algorithm, policy and contributing pages against plugrl-server found 21 discrepancies. This fixes them in both languages.
Wrong
plugrl-run-server --helplists policies only. Algorithms appear under<policy> <variant> --help.ppowas missing entirely, and now has a subsection: CleanRL's PPO, thedefaultanddppo-squarevariants, and which policies it takes.dppoextra.dppo-policyanddppo-gaussian-policydo. Without it they are silently absent from the CLI; the quoted log line did not exist.pi0-policydoes run underdppo, with theliberovariant, as E25 did.--policy.train-expert-only trueis a parse error. The pages now use the flag form.Stale. The training loop page now describes the server after #104 and #106:
discard_feedbackand are counted inserver/discarded_frames;global_stepcounts stored frames;on_policy = Falsefor replay buffers.The page also now covers resuming: the directory layout,
FileNotFoundErrorandFileExistsError, and the Ctrl-C checkpoint. It covers warm starts from another run's weights (--algo.policy-checkpoint-path,--algo.restore) as well.Missing
discard_feedback(an override must callsuper()),on_policy,derive_train_state,pre_learn/post_learn,init_optimizers, andload_checkpointon--resume. Registration needs both the config module and the class module imported.global_steps, and callsrecord_episode_metrics, without whichrollout/*stays 0.gaussian-policyanddppo-gaussian-policy; batch-first(B, H, ...)actions;action_dim/action_horizonfor the metadata; a runtime state that can be sliced per env; and thecriticDPPO and FPO require._predict_vanddt.InternalStateno longer exists; the pages name the real types.Checked. Every
plugrl-run-servercommand on these pages parses against the CLI.pi0-policywas checked with its OpenPI imports stubbed out, as OpenPI is not installed here. The custom algorithm template was loaded and each method called ondummy-policy.mkdocs build --strictpasses.Merge after plugrl-server #106 (
on_policy) and #107 (the dummy policy's defaults), which these pages describe.