《高级机器学习》:面向生成式人工智能时代的机器学习教材,涵盖强化学习、大模型训练与对齐、生成建模和多模态学习等内容。
-
Updated
Sep 21, 2026 - TeX
《高级机器学习》:面向生成式人工智能时代的机器学习教材,涵盖强化学习、大模型训练与对齐、生成建模和多模态学习等内容。
The implementation for our paper: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning.
Based on the virtual world built with IsaacSim, accomplish the training and deployment of real-world reinforcement learning for Hil-Serl.
Interactive Autonomous Navigation Research Lab for UGVs featuring localization, path planning, DWA, MPC, Adaptive MPC, reinforcement learning hooks, safety supervision, live visualization, benchmarking, replay, and automated report generation.
基于ms-swift的Qwen3_8B 金融推理两阶段后训练(LoRA SFT->GRPO)。
To associate your repository with the reforcement-learning topic, visit your repo's landing page and select "manage topics."