GPQA
Graduate-level Google-Proof Q&A benchmark — a hard-science and reasoning benchmark designed to resist internet-search-based cheating. GPQA Diamond (the hardest subset) is the primary discriminator between reasoning-mode frontier models in 2026.
Related terms
- MMLU-Pro
Massive Multitask Language Understanding — Pro version. A benchmark covering 57 subjects at professional/expert level. Successor to MMLU with reduced saturation on modern frontier models. Widely reported in model launches as a general-capability proxy.
- Reasoning Model
A class of LLMs (OpenAI o-series, Claude with extended thinking, DeepSeek-R1) trained to spend significant inference-time compute on internal reasoning before producing the final answer. Substantially better on math, coding, and multi-step problems at the cost of higher latency and token consumption.
Related on this site
Where this fits
GPQA is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.