mar. de 2025mai. de 2025ago. de 2025nov. de 2025fev. de 2026abr. de 2026jul. de 2026
README
A Survey of Reinforcement Learning for Large Reasoning Models
We welcome everyone to open an issue for any related work we haven’t discussed, and we’ll try to address it in the next release!
🎉 News
[2025-11-05] 🔥 Excited to release our paper list about Memory for Agents, covering breakthroughs in Context Management and Learning from Experience powering self-improving AI agents. Check it out: GitHub
[2025-10] 🎉 Honored to give talks at BAAI, Qingke Talk and Tencent Wiztalk! Here are the slides.
[2025-09-18] 🎉 We update the full list of papers in the category structure of the survey!
[2025-09-11] 🔥 Excited to release our RL for LRMs Survey! We’ll be updating the full list of papers in with a new category structure soon. Check it out: Paper.
[2025-08-15] 🔥 Introducing SSRL: an investigation for Agentic Search RL without reliance on external search engine. Check it out: GitHub and Paper.
[2025-05-27] 🔥 Introducing MARTI: A Framework for LLM-based Multi-Agent Reinforced Training and Inference. Check it out: Github.
[2025-04-23] 🔥 Introducing TTRL: an open-source solution for online RL on data without ground-truth labels, especially test data. Check it out: Github and Paper.
[2025-03-20] 🔥 We are excited to introduce collection of papers and projects on RL for reasoning models!
🎈 Citation
If you find this survey helpful, please cite our work:
@article{zhang2025survey,
title={A survey of reinforcement learning for large reasoning models},
author={Zhang, Kaiyan and Zuo, Yuxin and He, Bingxiang and Sun, Youbang and Liu, Runze and Jiang, Che and Fan, Yuchen and Tian, Kai and Jia, Guoli and Li, Pengfei and others},
journal={arXiv preprint arXiv:2509.08827},
year={2025}
}
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
2025-05
RFTF
RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
-
2025-05
VLA Generalization
What can rl bring to vla generalization? an empirical study
2025-02
ConRFT
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
2024-11
GRAPE
GRAPE: Generalizing Robot Policy via Preference Alignment
-
RLinf
RLinf: Reinforcement Learning Infrastructure for Agentic AI
-
-
EPO
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
Multi-Agent Systems
Date
Name
Title
Paper
Github
2025-10
AgentFlow
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
2025-09
SoftRankPO,
Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning
-
2025-09
BFS-Prover-V2
Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers
-
2025-08
MAGRPO
LLM Collaboration With Multi-Agent Reinforcement Learning
-
2025-06
AlphaEvolve
AlphaEvolve: A coding agent for scientific and algorithmic discovery
-
2025-06
JoyAgents-R1
JoyAgents-R1: Joint Evolution Dynamics for Versatile Multi-LLM Agents with Reinforcement Learning
-
2025-03
ReMA
ReMA: Learning to Meta-think for LLMs with Multi-agent Reinforcement Learning
2025-02
CTRL
Teaching Language Models to Critique via Reinforcement Learning
2025-02
Maporl
MAPoRL2: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning
2023-11
LLaMAC
Controlling large language model-based agents for large-scale decision-making: An actor-critic approach
-
Scientific Tasks
Date
Name
Title
Paper
Github
2025-09
Baichuan-M2
Baichuan-M2: Scaling Medical Capability with Large Verifier System
-
2025-08
CX-Mind
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
2025-08
MORE-CLEAR
MORE-CLEAR: Multimodal Offline Reinforcement learning for Clinical notes Leveraged Enhanced State Representation
-
2025-08
ARMed
Breaking Reward Collapse: Adaptive Reinforcement for Open-ended Medical Reasoning with Enhanced Semantic Discrimination
-
2025-08
ProMed
ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs
2025-08
OwkinZero
OwkinZero: Accelerating Biological Discovery with AI
-
2025-08
MolReasoner
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
2025-08
MedGR$^2$
MedGR$^2$: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning
-
2025-07
MedGround-R1
MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
2025-07
MedGemma
MedGemma Technical Report
-
2025-06
MMedAgent-RL
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
-
2025-06
Cell-o1
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
2025-06
MedAgentGym
MedAgentGym: Training LLM Agents for Code-Based Medical Reasoning at Scale
2025-06
Med-U1
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
2025-06
MedVIE
Efficient Medical VIE via Reinforcement Learning
-
2025-06
LA-CDM
Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
-
2025-06
ether0
Training a Scientific Reasoning Model for Chemistry
2025-06
Gazal-R1
Gazal-R1: Achieving State-of-the-Art Medical Reasoning with Parameter-Efficient Two-Stage Training
-
2025-05
DRG-Sapphire
Reinforcement Learning for Out-of-Distribution Reasoning in LLMs: An Empirical Study on Diagnosis-Related Group Coding
2025-05
BioReason
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
2025-05
EHRMIND
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
-
2025-04
Open-Medical-R1
Open-Medical-R1: How to Choose Data for RLVR Training at Medicine Domain
2025-04
ChestX-Reasoner
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
-
2025-04
BoxMed-RL
Reason Like a Radiologist: Chain-of-Thought and Reinforcement Learning for Verifiable Report Generation
-
2025-03
PPME
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience Learning
-
2025-03
DOLA
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization Agent
-
2025-02
Baichuan-M1
Baichuan-M1: Pushing the Medical Capability of Large Language Models
-
2025-02
MedVLM-R1
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
-
2025-02
Med-RLVR
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
-
2025-01
MedXpertQA
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
2024-12
HuatuoGPT-o1
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
-
Pro-1
Pro-1
-
rbio
rbio1 - training scientific reasoning LLMs with biological world models as soft verifiers
-
EPO
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
🌟 Acknowledgment
This survey is extended and refined from the original Awesome RL Reasoning Recipes repo. We are deeply grateful to all contributors for their efforts, and we sincerely thank for their all interest in Awesome RL Reasoning Recipes. The contents of the previous repository are available here.