Evaluation & Governance

Jailbreak

An adversarial input that bypasses a model's safety training to produce prohibited output. Common patterns: role-play framing ("pretend you are..."), hypothetical framing ("in a fiction..."), and multi-turn escalation. Related to but distinct from prompt injection.

Related terms

Where this fits

Jailbreak is part of the Evaluation & Governance vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.