Jailbreaking is using clever prompts to make an AI system ignore its safety constraints. Unlike prompt injection, where an attacker hides malicious instructions in external content, jailbreaking is usually direct. You’re in a chat with the AI, and you trick it into doing something it’s designed to refuse.
AI systems have guardrails. They’re trained to refuse harmful requests: “I can’t help with that,” “I won’t provide instructions for creating weapons,” “I don’t assist with illegal activities.” Those guardrails exist for good reasons. But guardrails aren’t unbreakable. And users, sometimes out of curiosity, sometimes with malicious intent, will find the cracks.
Why Guardrails Exist
Before we talk about breaking them, let’s understand why AI systems have safety constraints in the first place.



