It’s beginning to feel like AI companies are passing the baton when it comes to AI models breaking out of their testing environments, connecting to the internet, and compromising external organisations.

First, it was OpenAI earlier last week, when the company revealed that some of its AI models had escaped an internal testing environment and hacked into Hugging Face, an AI platform.

Now it’s Anthropic’s turn. The company has disclosed that three versions of its Claude AI accessed real organisations during internal security tests.

Anthropic Hires 32 AI Safety Experts to Fight Nuclear Risk
Anthropic will pay up to the high $200,000s for one job to think like someone trying to misuse Claude to build a weapon.

How the incident was uncovered

According to Anthropic, it launched a review of its own systems, analysing more than 141,006 evaluation runs. The review was prompted by the OpenAI-Hugging Face incident that surfaced days earlier.

During that review, Anthropic discovered three incidents in which Claude models had “gained unauthorized access to the production infrastructure of three different organizations.”

The earliest incident dates back to April. Two of the affected organisations had not detected the activity before Anthropic contacted them, while the company said it is still trying to reach the third.

How the models reached real systems

Subscribe for free to continue reading this article

Subscribe Subscribe

Already Have an Account? Log In