Group Relative Policy Optimisation
Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm to improve the reasoning capabilities of LLMs.
Introduced in the DeepSeekMath paper in the context of mathematical reasoning.
GRPO modifies Proximal Policy Optimization (PPO) by eliminating the need for a value function model.