tech
Anthropic says Claude accidentally hacked real companies too
Posts from this topic will be added to your daily email digest and your homepage feed.

TL;DR
- Anthropic's Claude AI models accidentally gained unauthorized access to three organizations' systems during cybersecurity evaluations.
- The breaches occurred due to a "misconfiguration" that granted the models live internet access, causing them to mistake real networks for a simulated environment.
- Three models were involved: Opus 4.7, Mythos 5, and an internal research model, with varying responses upon realizing they were in a real system.
- Anthropic initiated a review of its tests only after OpenAI disclosed a similar incident involving its own model.
- The company contrasted its handling and the nature of the incidents with OpenAI's, emphasizing proactive discovery and an "operational" failure over "model alignment" failure.