Loading the Elevenlabs Text to Speech AudioNative Player...

AI models can be dangerous; it's something many industry and policy experts have echoed for a while now. But now an incident with a government agency puts this concern into perspective.

On Tuesday, Britain’s AI Security Institute (AISI) disclosed that Claude Mythos went rogue during a cybersecurity evaluation, creating fake identities and attempting to insert malicious code into GitHub.

The agency described it as “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.”

What actually happened

AISI ran 122 cybersecurity challenge tests across multiple AI models. In 10 of those tests, AI agents took “autonomous, unsanctioned action on the live internet, targeting real people and organizations.”

In total, the agency identified 19 unsanctioned actions: 17 involving Anthropic’s Mythos 5 and two involving OpenAI’s GPT-5.6-Sol.

In the most severe case, an AI agent attempted to insert malicious code into an open-source GitHub project. To increase the chances of getting the code approved, it created fake online identities based on real people and used them to pressure the project’s maintainer into accepting the contribution.

Subscribe for free to continue reading this article

Subscribe Subscribe

Already Have an Account? Log In