LLM Reasoning
Photo by Hassan Pasha on Unsplash
Reasoning, in the context of LLMs, refers to the process of generating intermediate outputs before giving an answer. Typically, these outputs are tokens, which we call reasoning or thinking tokens. Reasoning in token-space is sometimes called Chain-of-Thought Reasoning (Raschka, 2026).
However, LLMs don't only reason in token-space. Latent Reasoning is where the LLM performs intermediate reasoning in hidden representations rather than generating a text token for each step (Hao et al., 2026).
The popularity of LLM reasoning seemed to explode after the release of the paper Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, in January 2022. The researchers showed that giving the model examples with intermediate reasoning steps substantially improved its ability on a range of arithmetic, commonsense and symbolic reasoning tasks (Wei et al., 2022), and it became a popular prompt engineering technique.
However, there had been papers exploring intermediate outputs in language models before that. In 2017, a DeepMind paper trained a model to generate step-by-step rationales before answering algebra word problems (Ling et al., 2017). In October 2021, OpenAI released the GSM8K dataset of maths word problems, fine-tuned GPT-3 on step-by-step solutions, and trained verifiers to pick the best of many candidate solutions (Cobbe et al., 2021). A month later, the Scratchpads paper showed that language models could handle multi-step tasks like long addition and executing code if they were trained, or even just prompted, to write their intermediate working to a "scratchpad" before the final answer (Nye et al., 2021).
What made chain-of-thought prompting different was that it needed no fine-tuning and worked across many kinds of tasks - it was an Emergent Property of sufficiently large LLMs, and a handful of worked examples in the prompt were enough to get a large model to show its working.
Shortly after the chain-of-thought paper, in March 2022, the Self-Consistency paper introduced a new decoding strategy for chain-of-thought prompting. Instead of greedily decoding a single chain of thought, the model samples several diverse reasoning paths and the most common final answer is selected, effectively a majority vote. The intuition is that a hard problem can often be solved in several different ways that all lead to the same correct answer. It improved chain-of-thought results on arithmetic and commonsense benchmarks, including a 17.9 percentage point gain on GSM8K (Wang et al., 2022).
Later that year, Large Language Models are Zero-Shot Reasoners (May 2022) showed that simply prompting the model with "Let's think step by step" before giving an answer could substantially improve its performance on a range of maths and logic tasks, without providing worked examples (Kojima et al., 2022).
A few years later, in late 2024, OpenAI released their o1 family of reasoning models, which were among the first examples of models trained with large-scale reinforcement learning to reason using chain of thought (OpenAI, 2024). The first releases, o1-preview and o1-mini, arrived in September 2024. The o1 API release in December 2024 added a reasoning-effort control, allowing developers to influence how much reasoning the model performs (openaiO1Model) (openaiAPIChangelog). o1 popularised the Test-Time Scaling paradigm, allowing the models to spend a configurable amount of computation when answering a question to improve the result.
In January 2025, DeepSeek released R1 and shared a training recipe, allowing the world to understand in more detail how a reasoning model can be trained. Their paper, DeepSeek-R1 Reasoning via Reinforcement Learning, distinguishes DeepSeek-R1-Zero, which applied reinforcement learning directly to a pretrained base model, from R1, which combined supervised fine-tuning and reinforcement learning. R1-Zero used rewards for correct answers and the required output format, allowing reasoning behaviour to emerge without an initial supervised fine-tuning stage (DeepSeek-AI et al., 2025) (original report, sections 2.2 and 2.3).
LLM reasoning is distinct from Agentic Reasoning, which extends the process through Tool Use, Planning, Reflection, Memory and so forth. Reasoning LLMs can be part of an agentic system, but generating reasoning tokens alone does not make a system agentic.
Reasoning in human cognition is also distinct. The exact nature of how we reason is not fully understood. Humans can generalise from a few examples and intuitively recognise abstract relationships, but it would be too strong to say that LLMs have none of these capabilities. Brown et al., 2020: Language Models are Few-Shot Learners demonstrated that models can learn tasks from a few examples in context (Brown et al., 2020). Generating reasoning tokens does not establish that a model reasons in the same way as a human. That said, LLMs can do more and more of what was once considered uniquely human.
Finally, there is Logical Reasoning, which allows us to arrive at conclusions from premises. More specifically, valid Deduction guarantees a true conclusion when its premises are true. Induction and abduction do not provide that guarantee. An LLM generating a chain of thought does not, by itself, guarantee a logically valid conclusion.
Reasoning, more generally, is about getting to an answer through some intermediate process.
References
Changelog \textbar OpenAI API. https://developers.openai.com/api/docs/changelog. ↩
O1 Model. https://developers.openai.com/api/docs/models/o1. ↩
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language Models are Few-Shot Learners. 2020. doi:10.48550/ARXIV.2005.14165. ↩
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. 2021. doi:10.48550/ARXIV.2110.14168. ↩
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training Large Language Models to Reason in a Continuous Latent Space. August 2026. arXiv:2412.06769, doi:10.48550/arXiv.2412.06769. ↩
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large Language Models are Zero-Shot Reasoners. 2022. doi:10.48550/ARXIV.2205.11916. ↩
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems. 2017. doi:10.48550/ARXIV.1705.04146. ↩
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. Show Your Work: Scratchpads for Intermediate Computation with Language Models. 2021. doi:10.48550/ARXIV.2112.00114. ↩
Sebastian Raschka. Build a Reasoning Model. Manning Publications, Erscheinungsort nicht ermittelbar, 2026. ISBN 978-1-63343-467-7. ↩
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. doi:10.48550/ARXIV.2203.11171. ↩
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. 2022. doi:10.48550/ARXIV.2201.11903. ↩
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jiawei Wang, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. Li, Shuang Zhou, Shaoqing Wu, Tao Yun, Tian Pei, Tianyu Sun, T. Wang, Wangding Zeng, Wanjia Zhao, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. 2025. doi:10.48550/ARXIV.2501.12948. ↩
OpenAI. OpenAI O1 System Card. 2024. ↩