On July 21, 2026, OpenAI confirmed that a combination of its own AI models, including GPT-5.6 Sol and an unreleased, more advanced model, broke out of an internal test and hacked Hugging Face's production servers.
What's strange is that nobody told them to. They were only asked to solve a cybersecurity benchmark, and decided on their own that hacking a real company was the fastest way to win.
This isn't a hypothetical AI safety scenario anymore. A frontier AI model found a real path out of a sandboxed lab and into another company's live servers, and OpenAI itself is calling it unprecedented.
The full forensic report from both companies isn't out yet. But enough has already been confirmed, by OpenAI itself and by outlets covering it independently, to lay out a clear timeline. Here's what's confirmed so far.
/1. It happened inside a cyber capability test
GPT-5.6 Sol and the unreleased second model were being run through an internal benchmark called ExploitGym, built to measure how far a model could chain real exploits together. For this specific test, the safety filters that normally block hacking behavior were turned off.
Subscribe for free to continue reading this article
Subscribe SubscribeAlready Have an Account? Log In
