Skip to content

Anthropic - Our models can breach containment as well

The door was unlocked, but Claude pushed it open...

Anthropic - Our models can breach containment as well
Image credit: https://unsplash.com/@katjaano

A week after OpenAI said an agent hacked Hugging Face during testing, Anthropic says its own Claude models have also breached three businesses after mistakenly being given internet access.

The AI lab said it had looked back through testing transcripts following OpenAI’s 21 July disclosure and found three models, including Mythos, had hacked companies during capture-the-flag challenges since April.

But while OpenAI claimed its agent hacked its way to the internet, Anthropic blamed evaluation firm Irregular for giving internet access to sandboxes that should have been closed off.

In a blog post shared Thursday, it said: “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise… [and] compromised the impacted organizations’ infrastructure using basic techniques.

This content is for members only

Subscribe
Add The Stack on Google