Home /paper

DeepSeek-R1 Reasoning via Reinforcement Learning

Notes on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning by DeepSeek-AI.

DeepSeek train a model, DeepSeek-R1-Zero, using large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, and find it demonstrates remarkable reasoning capabilities (DeepSeek-AI et al., 2025).

However, R1-Zero has problems like poor readability and language mixing. So they build DeepSeek-R1, which adds multi-stage training and cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1-1217 on reasoning tasks.

They open-source:

  • DeepSeek-R1-Zero
  • DeepSeek-R1
  • six dense models (1.5B, 7B, 8B, 14B, 32B, 70B) distilled from DeepSeek-R1, based on Qwen and Llama

The usual steps for creating an LLM (top), compared with DeepSeek-R1-Zero, which goes straight from the base model to reinforcement learning, and DeepSeek-R1, which alternates supervised fine-tuning and RL.

DeepSeek-R1-Zero

They use DeepSeek-V3-Base as the base model and train it with Group Relative Policy Optimisation. Pure reinforcement learning, no supervised fine-tuning. Incredible.

The rewards are simple and rule-based: an accuracy reward for getting the right answer, and a format reward for putting the reasoning between <think> tags (section 2.2.2). Nobody shows the model how to reason.

  • Achieved 71.0% accuracy on AIME 2024, up from 15.6% at the start of training
  • Learned to generate longer chains of thought as training went on
  • Developed self-verification and reflection, including an "aha moment" where it learns to stop and re-evaluate its approach

However, the model had trouble with readability and language mixing, which led to its successor.

DeepSeek-R1

DeepSeek-R1 uses a multi-stage training pipeline:

  1. Cold start: fine-tune the base model on thousands of long chain-of-thought examples
  2. Reasoning-oriented RL, with an extra reward for language consistency
  3. Rejection sampling to create new SFT data, then fine-tune again
  4. A second round of RL covering all scenarios, including helpfulness and harmlessness

DeepSeek-R1 matches or exceeds OpenAI-o1-1217 across multiple benchmarks:

  • 79.8% on AIME 2024
  • 97.3% on MATH-500
  • 90.8% on MMLU
  • 96.3 percentile on Codeforces

Distillation

They also show the reasoning can be distilled into smaller models, by fine-tuning them on data generated by DeepSeek-R1:

  • DeepSeek-R1-Distill-Qwen-7B achieves 55.5% on AIME 2024
  • The 32B variant reaches 72.6%

Interestingly, distilling from R1 worked better than running large-scale RL on the smaller model directly.

What didn't work

They also share their unsuccessful attempts:

  • Process Reward Models (PRMs) were hard to scale and prone to reward hacking
  • Monte Carlo Tree Search (MCTS) struggled with the huge search space of token generation

Future work

  • Improving general capabilities, like function calling and multi-turn interactions
  • Fixing language mixing
  • Reducing sensitivity to prompts (few-shot prompting made it worse)
  • Improving software engineering performance

Takeaways

This paper shows that reinforcement learning can be the main driver of reasoning in language models, reducing the need for large supervised datasets. And because DeepSeek open-sourced the models and shared the recipe, we can finally see how a reasoning model is trained.

References

DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. Li, Shuang Zhou, Shaoqing Wu, Tao Yun, Tian Pei, Tianyu Sun, T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. doi:10.48550/ARXIV.2501.12948. ↩