A Conceptual Taxonomy of Jailbreak Techniques
Explore different jailbreak methods targeted at LLMs and agentic systems, focusing on how attackers manipulate instructions, roles, and constraints to circumvent security. Learn to classify these techniques by their mechanisms and evaluate claims within threat models to understand vulnerabilities and strengthen system defenses.
A jailbreak is any attempt to circumvent an application’s intended constraints on an LLM’s behavior, usually to break integrity goals like “follow the policy” or confidentiality goals like “keep secrets”. In a customer-support assistant, those constraints often live in a system prompt, policy text, and surrounding application logic, even when the user only sees a chat box.
Three attack intents can look similar on the surface but work through different mechanisms. One tries to reinterpret the rules so the policy seems inapplicable, another tries to change the assistant’s role or objective so it prioritizes the attacker’s outcome, and a third hides intent through indirection so the model fails to recognize what it is doing. A taxonomy helps you discuss papers and incidents by mechanism and assumption, instead of sharing exploit strings that are easy to copy. This lesson sticks to patterns, not payloads, and stays at the level of conceptual levers and threat models. ...