Emergent Abilities
Emergent Abilities, in the context of Large Language Models, are abilities that are "not present in smaller models but [are] present in larger models" (Wei et al., 2022). Performance stays near random as models get bigger, then jumps once they pass some scale. You can't predict it by extrapolating from the smaller models.
The term comes from the idea of Emergence more generally. Wei et al. root it in physicist Philip Anderson's 1972 essay "More Is Different" (Anderson, 1972), summarising it as: "Emergence is when quantitative changes in a system result in qualitative changes in behavior". In LLMs, the quantity is scale: training compute and parameter count.
Chain-of-Thought Prompting is one of their main examples. It only beats standard prompting at around 100B parameters (Wei et al., 2022). Scratchpads and Self-Consistency decoding are also on their list.
However, not everyone agrees that these jumps are real. Schaeffer et al. argue that emergent abilities "appear due to the researcher's choice of metric rather than due to fundamental changes in model behavior with scale" (Schaeffer et al., 2023). All-or-nothing metrics like exact match make performance look like it jumps suddenly. Measure the same models with a smoother metric and the improvement is gradual and predictable.
References
P. W. Anderson. More Is Different. Science, 177(4047):393–396, 1972. doi:10.1126/science.177.4047.393. ↩
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are Emergent Abilities of Large Language Models a Mirage? 2023. arXiv:2304.15004, doi:10.48550/arXiv.2304.15004. ↩
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. Emergent Abilities of Large Language Models. 2022. doi:10.48550/ARXIV.2206.07682. ↩
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. 2022. doi:10.48550/ARXIV.2201.11903. ↩