LLM Training Pipeline
Photo by Taylor Vick on Unsplash
Large Language Models training is typically structured into two phases: pre-training and post-training (Raschka, 2026).
Pre-training
In the Pre-training step, LLMs are trained on giant collections of data, often gathered from the internet, books and other sources. These collections can reach terabytes or petabytes, depending on the training run. Autoregressive LLMs are typically trained with the next-token-prediction objective. To get good at doing that, the model needs to learn patterns in language and the relationships represented in the dataset, developing a sort of "understanding" of it. Large pre-training runs can take weeks to months and are typically extremely expensive (Raschka, 2026).
Post-training
Once we have a pre-trained LLM, we fine-tune it for a task, like responding to chat messages. Techniques include Supervised Fine-Tuning, including instruction tuning and preference tuning to teach it to respond to user queries in ways people prefer. Instruction tuning is a form of supervised fine-tuning, rather than a name for all supervised fine-tuning (Raschka, 2026) (Ouyang et al., 2022).
Reinforcement Learning from Human Feedback (RLHF) is a form of preference tuning. Direct Preference Optimisation (DPO) is another approach: it learns directly from preferred and rejected responses without a separate reward-model training stage or an RL rollout loop (Ouyang et al., 2022) (Rafailov et al., 2023).
References
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. March 2022. arXiv:2203.02155, doi:10.48550/arXiv.2203.02155. ↩ 1 2
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. 2023. doi:10.48550/ARXIV.2305.18290. ↩
Sebastian Raschka. Build a Reasoning Model. Manning Publications, Erscheinungsort nicht ermittelbar, 2026. ISBN 978-1-63343-467-7. ↩ 1 2 3