Evaluation & Governance

Eval Harness

The full pipeline that runs a set of evaluation prompts through one or more models, scores the outputs (via reference match, LLM-as-judge, or human review), and reports results. Common open-source harnesses: lm-eval-harness, HELM, OpenAI evals. Enterprise programs usually build a custom harness against a golden-dataset backend.

Related terms

Where this fits

Eval Harness is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.