I’ve always wanted to see how far most AI chatbots are willing to go, whether their safety guardrails actually hold, and, if they do, just how strong they are.

I remember the controversy surrounding Google's first chatbot, then called Bard, back in early 2023, when it produced some unhinged answers. It made me wonder how much AI safety had improved since then.

Over the past two weeks, I tested ChatGPT, Claude, DeepSeek, Gemini, and Grok using the same categories of jailbreak techniques to see which models consistently enforced their safety policies and which were easier to bypass.

To keep the comparison as fair as possible, I used broadly comparable prompts across each model, including language-based prompts, role-playing, hypothetical scenarios, creative-writing framing, prompt obfuscation, and repeated attempts to reframe rejected requests.

When one approach failed, I changed the phrasing rather than the objective. This wasn't a scientific benchmark or penetration test, but a practical experiment designed to compare how consistently each chatbot responded to adversarial prompting.

And the results surprised me. ChatGPT and Claude consistently resisted every meaningful attempt I made to bypass their guardrails. The other models varied considerably.

I Tested 5 AI Coding Tools, and Only Two Didn’t Fail
One of these five tools let me publish a page that leaked every guest’s name and email: Lovable, Replit, Antigravity, Cursor, and Claude Code.

ChatGPT’s safety guardrails stood tall against everything I threw at it

Subscribe for free to continue reading this article

Subscribe Subscribe

Already Have an Account? Log In