MMLU-Pro
Massive Multitask Language Understanding — Pro version. A benchmark covering 57 subjects at professional/expert level. Successor to MMLU with reduced saturation on modern frontier models. Widely reported in model launches as a general-capability proxy.
Related terms
- GPQA
Graduate-level Google-Proof Q&A benchmark — a hard-science and reasoning benchmark designed to resist internet-search-based cheating. GPQA Diamond (the hardest subset) is the primary discriminator between reasoning-mode frontier models in 2026.
- Eval Harness
The full pipeline that runs a set of evaluation prompts through one or more models, scores the outputs (via reference match, LLM-as-judge, or human review), and reports results. Common open-source harnesses: lm-eval-harness, HELM, OpenAI evals. Enterprise programs usually build a custom harness against a golden-dataset backend.
Related on this site
Where this fits
MMLU-Pro is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.