Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/media/coverage/coverage.json

Large diffs are not rendered by default.

Binary file modified docs/media/coverage/pi0-fpo.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file modified docs/media/coverage/pi0-fpo.mp4
Binary file not shown.
19 changes: 14 additions & 5 deletions docs/vla.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,14 +3,17 @@
The home page shows a full-size pi0.5 running across PlugRL's boundary, with
its evaluations matching openpi's. This page is the rest of that record: what
fine-tuning it with reinforcement learning through PlugRL has done so far.
In short, it has not made the policy better.
In short, it has not made the policy better. The collapses the first runs
showed were defects of ours, and with them fixed FPO leaves the policy about
where it started.

<div class="cov" data-part="vla" data-src="/media/coverage/coverage.json"></div>

All three clips start from the same scene, the first one the released policy
solves, and the numbers under them come from fifty-episode evaluations. Over
seven such evaluations the released policy scored between 28 and 37. Click a
clip to see the two commands behind it.
seven such evaluations the released policy scored between 28 and 37. Where an
experiment ran two seeds, the clip and the number are the lower-scoring
seed's. Click a clip to see the two commands behind it.

## How it went

Expand All @@ -32,8 +35,14 @@ than FPO++'s. Ten DPPO iterations left it at 20, three standard deviations
below the released policy's mean
([E36](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e36-pi0-longer)).

A run with FPO++'s fine-tuning in full is under way, to check whether what
remains of the collapse is still ours.
That fall was ours too. With the rest of FPO++'s fine-tuning in place (its
optimizer, gradient clipping, critic learning rate, raw rewards, lambda and
clip), ten FPO iterations left the same two seeds at 40 and 27 of 50
([E42](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e42-pi0-fpo-plus-plus)).
Neither reaches the bar set before the run, 42 of 50, so this is FPO holding
pi0.5, not improving it. The lower seed, whose clip is above, was drifting
the way E36's runs fell, only far more slowly. Whether it holds past ten
iterations is open.

The scripts that recorded the clips are in
[figures/coverage](https://github.com/PlugRL/plugrl-server/tree/main/figures/coverage).
13 changes: 9 additions & 4 deletions docs/vla.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,13 +2,14 @@

主页展示了一个完整尺寸的 pi0.5 穿过 PlugRL 的边界跑通,评估结果和 openpi 公布的一致。
这一页是那份记录的其余部分:通过 PlugRL 用强化学习微调它,到目前为止做到了什么。
简单说,还没能让策略变好。
简单说,还没能让策略变好。前几次跑出来的崩塌都是我们自己的缺陷,修掉之后 FPO 基本让策略
停在原来的水平。

<div class="cov" data-part="vla" data-src="/media/coverage/coverage.json"></div>

三段视频都从同一个场景开始,也就是原版策略能完成的第一个场景;视频下面的数字来自 50 个
回合的评估。原版策略在七次这样的评估里得分在 28 到 37 之间。点一下视频,会显示它背后的
两条命令。
回合的评估。原版策略在七次这样的评估里得分在 28 到 37 之间。实验跑了两个种子的,视频和
数字都取分数较低的那个。点一下视频,会显示它背后的两条命令。

## 经过

Expand All @@ -26,7 +27,11 @@
还是自己的默认设置,没换成 FPO++ 的;十轮 DPPO 后是 20/50,比原版策略的平均低了三个
标准差([E36](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e36-pi0-longer))。

换上 FPO++ 完整设置的实验正在跑,用来确认剩下的崩塌还是不是我们的问题。
这次下跌也是我们的问题。把 FPO++ 其余的微调设置补齐(优化器、梯度裁剪、critic 学习率、
原始奖励、lambda 和 clip)之后,同样两个种子跑十轮 FPO,分别是 40/50 和 27/50
([E42](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e42-pi0-fpo-plus-plus))。
两个都没达到跑之前定好的 42/50,所以这是 FPO 稳住了 pi0.5,不是把它变好了。分数较低的
那个种子(上面的视频就是它)在往 E36 那样的方向漂,只是慢得多;再往后跑还稳不稳,还不知道。

录这些视频的脚本在
[figures/coverage](https://github.com/PlugRL/plugrl-server/tree/main/figures/coverage)。
Loading