Chain-of-Thought Reasoning
Chain-of-Thought Reasoning is an LLM Reasoning technique where an LLM can reason in token space: in other words, it generates reasoning traces as text before generating an answer.
It can be elicited through Chain-of-Thought Prompting, where the model is given examples with intermediate reasoning, or through Think Step-by-Step prompting without examples. Models like OpenAI's o1 and DeepSeek-R1-Zero were subsequently trained to reason using reinforcement learning. For R1-Zero, the rewards evaluated answer correctness and output format, rather than prescribing each reasoning step.
See LLM Reasoning for the history and its relationship to Agentic Reasoning, human cognition and Logical Reasoning.