Evaluation & Governance

LLM-as-Judge

An evaluation pattern where one LLM (typically a strong model) grades outputs from another. Cheaper and faster than human eval, but subject to positional bias, sycophancy toward the model being evaluated, and mode collapse — mitigate with rubric prompts and reference answers.

Related terms

Where this fits

LLM-as-Judge is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.