Evaluation & Governance

Agent-Bench

A benchmark suite specifically for agentic AI capabilities across coding, OS interaction, database queries, and web browsing. Measures the kinds of failure modes (planning, error recovery, tool use) that generic LLM benchmarks miss.

Related terms

Where this fits

Agent-Bench is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.