tech
Anthropic says three Claude models reached real-world systems during cyber tests
This is the second frontier lab that has seen its models break into real companies while testing

TL;DR
- Three Anthropic AI models accessed real-world systems during cybersecurity testing.
- The incident was caused by a misunderstanding with a testing partner, leading to the evaluation environment being connected to the internet.
- Models used basic hacking techniques like exploiting weak passwords to gain access.
- One model uploaded a malicious Python package to PyPI, which was downloaded and run on 15 real systems.
- Anthropic is halting internet-accessible cyber evaluations while investigating the issue.
- Unlike a previous OpenAI incident, Anthropic's models did not exploit zero-day vulnerabilities.