LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
-
Updated
Sep 7, 2026 - Python
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
Codebase of GRPO: Implementations and Resources of GRPO and Its Variants
Open Ended Medical Reinforcement Learning
Exact finite-group identity behind GRPO reward standardization, unifying GRPO / Dr. GRPO / DAPO for RLVR and LLM reasoning. Paper + code.
Process-supervised RL for a multi-step reasoning agent — DAPO + a learned Process Reward Model (PRM) training a Qwen3-8B Planner. A modern, from-scratch rebuild of the AgentFlow paper (ICLR 2026).
Sync vs fully-async agentic RL on verl: multi-turn GRPO, long-tail rollout profiling, staleness ablations — quantifying when async pays off.
Unofficial PyTorch reproduction for DAPO: An Open-Source LLM Reinforcement Learning System at Scale.
To associate your repository with the dapo topic, visit your repo's landing page and select "manage topics."