AI Fundamentals

Interpretability

The subfield concerned with understanding what happens inside AI models — which features they represent, how they compose, why they fire on given inputs. Mechanistic interpretability is a leading research direction; sparse autoencoders are one recent technique.

Related terms

Next steps

Where this fits

Interpretability is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.