Agentic AI

Agent Evals

Systematic testing of AI agents against representative tasks — measuring task completion, tool-use correctness, cost, latency, and safety across many runs. Distinct from single-turn LLM evals because agent behavior spans multiple steps and depends on tool outputs, memory, and planning quality. Considered table stakes before shipping agents to production.

Related terms

For developers

Call this framework and its tools from your own agent via the Model Context Protocol (MCP) server. Works with Claude Desktop, Cursor, Zed, Continue, and the OpenAI Agents SDK.

Where this fits

Agent Evals is part of the Agentic AI vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.