AI Fundamentals

GRPO

Group Relative Policy Optimization — a reinforcement-learning alignment technique used in DeepSeek-R1 and other reasoning-model training. Compares groups of sampled outputs against each other to compute advantages, avoiding the value-model overhead of PPO.

Related terms

Next steps

Where this fits

GRPO is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.