Reinforcement Learning from Human Feedback
Reinforcement Learning from Human Feedback (RLHF) is a form of preference tuning used during post-training in the LLM Training Pipeline (Raschka, 2026).
Start with a pre-trained language model, trained on a language modelling objective like next-token prediction or masked-token prediction.
The model can first undergo Supervised Fine-Tuning on human-labelled demonstration data. InstructGPT used this sequence before reward-model training and RL fine-tuning (Ouyang et al., 2022).
Reward Model Training
A reward model is trained from human preference data (e.g. pairwise comparisons between model outputs) (Ouyang et al., 2022).
RL Fine-Tuning (e.g. Proximal Policy Optimization (PPO))
The language model is fine-tuned using reinforcement learning, with the reward model as feedback (Ouyang et al., 2022).
The RL loop samples responses from the LLM, scores them using the reward model, and updates the LLM weights accordingly. Pairwise comparisons are used to train the reward model; the RL loop does not need to sample responses in pairs (Ouyang et al., 2022).
Human feedback can be collected ahead of time. That does not make PPO training offline RL: the policy generates fresh responses during training, and the reward model scores them without needing a human to review each response (Ouyang et al., 2022).
References
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. March 2022. arXiv:2203.02155, doi:10.48550/arXiv.2203.02155. ↩ 1 2 3 4 5
Sebastian Raschka. Build a Reasoning Model. Manning Publications, Erscheinungsort nicht ermittelbar, 2026. ISBN 978-1-63343-467-7. ↩