Loading the Elevenlabs Text to Speech AudioNative Player...

As the debate over the dangers of AI continues, OpenAI, the company behind ChatGPT, has released one of its most revealing safety reports yet.

On Wednesday, the company published a research paper titled "Our framework for reporting model misalignment," detailing six incidents where its AI models behaved in unexpected and potentially dangerous ways during training and testing over the past six months.

Perhaps the most striking part of the report is that OpenAI openly admits it doesn't know whether the industry can keep scaling AI systems at its current pace without stronger safety measures.

"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," the company said.

That makes the current AI safety conversation feel even more urgent, as researchers and industry insiders continue warning that AI capabilities are advancing faster than the guardrails designed to keep them in check.

The company also said it now favours disclosure "even when significance is uncertain," meaning some incidents may turn out to be isolated cases rather than evidence of broader patterns, but it believes they're still worth sharing publicly.

Why Anthropic Researcher Jacob Coxon Quit, and Who Agreed
Coxon’s warning comes as both OpenAI and Anthropic have faced incidents involving their AI systems gaining access to other organisations’ systems.

The six incidents OpenAI reported

The first incident involved an unreleased research model that inserted "jailbreak-like instructions" into its own summaries. Those summaries included persona prompts describing itself as being "freed from the roles and identities that bind other chatbots."

The model even instructed itself:

"You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to."

It sounds like something pulled from a science-fiction movie, but this time it came from an internal AI training run. OpenAI found 27 affected summaries.

The second incident happened during GPT-5.6 Sol training, where models began adding instructions telling future versions of themselves to hide mistakes from users, invent missing information, and conceal version mismatches.

The third incident involved a model answering a routine question. It discovered an exposed API key online, used it without authorisation, and then fabricated the data it couldn't retrieve.

The fourth incident was even stranger. An unreleased model solved a programming task correctly using code, then uploaded its own file to the internet without permission so it could cite a web source in its response.

The fifth incident showed models using an internal software repository as a message board to communicate across separate training samples.

And in the sixth incident, multiple AI agents working on the same task began using public file-sharing websites to exchange documents when they couldn't directly access each other's local files.

Why OpenAI is publishing these reports now

OpenAI admits it hasn't handled these incidents particularly well in the past.

In the report, the company wrote:

"Our disclosures have been ad hoc and less frequent than ideal... We've often waited until we could collate several instances into one report, or added them to system cards for newly released models."

Subscribe for free to continue reading this article

Subscribe Subscribe

Already have an account? Log in