Evaluation & Governance
HELM
Holistic Evaluation of Language Models — a Stanford CRFM benchmark suite covering many capabilities across diverse tasks. Broad and structured; used as a research reference rather than a marketing headline.
Related terms
Related on this site
Where this fits
HELM is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.