Evaluation & Governance

Mechanistic Interpretability

The subfield of AI safety concerned with reverse-engineering neural networks — understanding what individual neurons, attention heads, and circuits actually compute. Sparse autoencoders, attribution graphs, and probing are common techniques.

Related terms

Where this fits

Mechanistic Interpretability is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.