Evaluation & Governance

SWE-Bench Verified

A benchmark of real GitHub issues from popular Python repos where the model must produce a passing patch. The "Verified" subset filters for well-specified problems. In 2026 the leading indicator of coding-agent capability, with frontier models scoring 60-75.

Related terms

Where this fits

SWE-Bench Verified is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.