RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.
-
Updated
Jul 1, 2026 - Python
RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.
Reinforcement learning algorithms with mathematical derivations and Sutton & Barto figure reproductions.
基于 Sutton & Barto《强化学习》第2版的系统性精读笔记。MDP、贝尔曼方程、Q-learning、策略梯度 —— 专业排版 + 公式推导 + 直觉解释。
🎯 Reinforcement Learning portfolio · 5 envs · 20+ algorithms (DP, MC, TD, SARSA, Q-Learning, REINFORCE, MCTS) + Deep RL (DQN, AlphaZero) · 71 pytest tests
Sutton & Barto'nun Reinforcement Learning: An Introduction kitabi uzerine ozgun Turkce calisma notlari + mobil uyumlu web okuyucu
Reinforcement learning implementations and experiments based on Sutton & Barto, including Blackjack, Gambler's Problem, and Multi-Armed Bandits.
Deep Reinforcement Learning from mathematical foundations to PyTorch implementations: rigorous proofs, step-by-step derivations of objectives and gradient estimators, and reproducible experiments on Gymnasium and MuJoCo benchmarks spanning policy gradients, actor-critic methods, value-based learning, and continuous control.
Every algorithm from Sutton & Barto Ch 1-13 - tabular RL in NumPy, deep RL in PyTorch
Sutton and Barto's 10-armed bandit testbed with epsilon-greedy, optimistic initialization, and UCB, averaged over 2,000 problems.
Multi armed bandits code based on Sutton & Barto Reinforcement Learning Book
Learn reinforcement learning by climbing one ladder of algorithms on a gridworld you own — every learner scored against the exact optimal values from dynamic programming.
Reinforcement Learning: Gymnasium/PettingZoo (Farama), Sutton & Barto, HuggingFace RL — session material for the DS/ML Collaborative Learning Group
Fork of ShangtongZhang/reinforcement-learning-an-introduction - Python implementations of algorithms from Sutton and Barto's RL textbook (2nd Edition)
Interactive RL learning platform: 13 chapters from Sutton & Barto, 18K+ lines, fill-in-the-blank exercises with bilingual explanations. Bandits → DP → MC → TD → Policy Gradient → DQN → PPO → SAC → MARL → RLHF
n-armed bandit algorithms comparison + simulation app
Windy Gridworld with Q-learning, SARSA, and Expected SARSA, including learning curves and the trained agent's path.
To associate your repository with the sutton-barto topic, visit your repo's landing page and select "manage topics."