Home /permanent

Group Relative Policy Optimisation

Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm to improve the reasoning capabilities of LLMs.

Introduced in the DeepSeekMath paper in the context of mathematical reasoning.

GRPO modifies Proximal Policy Optimization (PPO) by eliminating the need for a value function model.