Self-Consistency Improves Chain of Thought Reasoning in Language Models
Notes on Self-Consistency Improves Chain of Thought Reasoning in Language Models by Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou.
Self-Consistency is a decoding strategy for Chain-of-Thought Prompting, introduced by Google in March 2022, shortly after the original chain-of-thought paper (Wang et al., 2022).
Instead of greedily decoding a single chain of thought, you sample several diverse reasoning paths from the model, then pick the final answer that comes up most often: effectively a majority vote. The intuition is that a hard problem can usually be solved in several different ways that all lead to the same correct answer, while wrong reasoning tends to lead to different wrong answers.
It needs no extra training, fine-tuning or verifier. It boosted chain-of-thought results on arithmetic and commonsense benchmarks, including GSM8K (+17.9%), SVAMP (+11.0%), AQuA (+12.2%), StrategyQA (+6.4%) and ARC-challenge (+3.9%).
The trade-off is compute: you pay for every sampled path. And a majority vote only works when answers can be compared directly, like a number or a multiple-choice option, not open-ended text.
It's one of the earliest examples of Test-Time Scaling: spend more compute at inference time to get a better answer. See LLM Reasoning for where it fits in the history.
References
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. doi:10.48550/ARXIV.2203.11171. ↩