Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 6 additions & 3 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,8 +84,9 @@ confirm the two sides talk to each other, not to train anything.

PlugRL keeps the policy, the algorithm and the environment apart, so the
question worth answering is which combinations actually work. Below is every
combination of the two MLP policies and the two algorithms on four tasks, and
pi0.5 on LIBERO. The border and its label say what the experiments found. The
combination of the two MLP policies and the two algorithms on four tasks; the
baseline they are measured against, a Gaussian policy with PPO run as CleanRL
runs it; and pi0.5 on LIBERO. The border and its label say what the experiments found. The
line under each clip is the training return of all three seeds, drawn on one
scale per column, so a flat line really is flat. Hover over a cell to play
it; click or tap it to play it and see the two commands that trained it.
Expand All @@ -98,7 +99,9 @@ Each clip comes from the final checkpoint of the seed whose last ten
iterations were the median of three. That checkpoint was evaluated for five
episodes, and the clip is the episode with the median return, not the best
one. `dppo-policy · FPO` is missing because it cannot exist: FPO trains flow
policies only. The three pi0.5 clips all start from the same scene, the first
policies only. `gaussian-policy · PPO` has no square cell because it was not
run there; on the three MuJoCo tasks it ends about where CleanRL's own runs
do ([E38](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e38-gaussian-ppo)). The three pi0.5 clips all start from the same scene, the first
one the released policy solves, and the numbers under them come from
fifty-episode evaluations. pi0.5 used to fall to zero after one FPO
iteration; the defect was ours, in how our FPO scored an action chunk, and
Expand Down
6 changes: 4 additions & 2 deletions docs/index.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,8 @@ dummy 算法的 `learn` 是一个 sleep,不会移动任何权重。它用来
## 在它上面能跑什么

PlugRL 把策略、算法和环境拆开,所以真正值得回答的问题是:哪些组合真的能跑。下面是
两个 MLP 策略和两个算法在四个任务上的全部组合,外加 pi0.5 在 LIBERO 上的情况。边框
两个 MLP 策略和两个算法在四个任务上的全部组合,加上作为对照基线、照 CleanRL 原样跑的高斯策略 + PPO,
以及 pi0.5 在 LIBERO 上的情况。边框
和上面的标签写的是实验得出的结论。每段视频下面那条线是三个种子的训练回报,同一列用同
一个纵轴,所以平的线就是真的没动。鼠标悬停就能播放;点一下格子,除了播放,还会在下面显示训练
它的那两条命令。从一格换到另一格,变的只有指定策略、算法和任务的那几个词。
Expand All @@ -84,7 +85,8 @@ PlugRL 把策略、算法和环境拆开,所以真正值得回答的问题是

每段视频都取自三个种子里最后十轮回报居中的那个种子的最终检查点。这个检查点评估了五
个回合,放出来的是回报居中的那一回合,不是最好的那一回合。表里没有 `dppo-policy · FPO`,
因为这个组合不存在:FPO 只能训练流策略。pi0.5 的三段视频都从同一个场景开始,也就是原
因为这个组合不存在:FPO 只能训练流策略。`gaussian-policy · PPO` 在方块任务那格是空的,因为没在那里跑过;
它在三个 MuJoCo 任务上的回报和 CleanRL 自己报告的差不多([E38](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e38-gaussian-ppo))。pi0.5 的三段视频都从同一个场景开始,也就是原
版策略能完成的第一个场景;视频下面的数字来自 50 个回合的评估。pi0.5 以前一轮 FPO 后
就掉到零,问题出在我们这边——这里的 FPO 给动作块打分的方式不对,换成 FPO++ 的打分方式后
策略就不再被毁([E32](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e32-pi0-fpo-plus-plus))。生成这些内容的脚本在
Expand Down
Binary file modified docs/media/coverage-grid.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
2 changes: 1 addition & 1 deletion docs/media/coverage/coverage.json

Large diffs are not rendered by default.

Binary file added docs/media/coverage/gaussian-ppo-cheetah.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/media/coverage/gaussian-ppo-cheetah.mp4
Binary file not shown.
Binary file added docs/media/coverage/gaussian-ppo-hopper.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/media/coverage/gaussian-ppo-hopper.mp4
Binary file not shown.
Binary file added docs/media/coverage/gaussian-ppo-walker.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/media/coverage/gaussian-ppo-walker.mp4
Binary file not shown.
Loading