dapo
Here are 16 public repositories matching this topic...
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
-
Updated
Sep 7, 2026 - Python
Codebase of GRPO: Implementations and Resources of GRPO and Its Variants
-
Updated
Dec 6, 2025 - Python
Open Ended Medical Reinforcement Learning
-
Updated
Sep 12, 2026 - Python
Unified reproductions, benchmarking & evaluation for perception-aware RLVR in VLMs. 11 methods · 25+ benchmarks · Official code of CGPO (ACM MM 2026 Oral).
-
Updated
Oct 4, 2026 - Python
An end-to-end framework for Agentic Reinforcement Learning.
-
Updated
Sep 28, 2026 - Python
Exact finite-group identity behind GRPO reward standardization, unifying GRPO / Dr. GRPO / DAPO for RLVR and LLM reasoning. Paper + code.
-
Updated
Jul 2, 2026 - TeX
Sync vs fully-async agentic RL on verl: multi-turn GRPO, long-tail rollout profiling, staleness ablations — quantifying when async pays off.
-
Updated
Sep 8, 2026 - Python
Unofficial PyTorch reproduction for DAPO: An Open-Source LLM Reinforcement Learning System at Scale.
-
Updated
Jul 2, 2026 - Python
让 RL 工程里的坑不再只活在某个人的经验里。**当前是一个 demo ——** 只给结构、不给内容:36 个纯文本 Markdown 节点 + 一个只查自洽、不判对错的校验器 + 一个离线确定性检索器(无模型)。 | Make RL engineering pitfalls survive beyond one person's memory. **A demo —** structure, not content: 36 plain-text Markdown nodes, a consistency-only validator, and an offline deterministic matcher (no model).
-
Updated
Sep 30, 2026 - HTML
Add this topic to your repo
To associate your repository with the dapo topic, visit your repo's landing page and select "manage topics."