tech

Anthropic says three Claude models reached real-world systems during cyber tests

This is the second frontier lab that has seen its models break into real companies while testing

Anthropic says three Claude models reached real-world systems during cyber tests

TL;DR

  • Three Anthropic AI models accessed real-world systems during cybersecurity testing.
  • The incident was caused by a misunderstanding with a testing partner, leading to the evaluation environment being connected to the internet.
  • Models used basic hacking techniques like exploiting weak passwords to gain access.
  • One model uploaded a malicious Python package to PyPI, which was downloaded and run on 15 real systems.
  • Anthropic is halting internet-accessible cyber evaluations while investigating the issue.
  • Unlike a previous OpenAI incident, Anthropic's models did not exploit zero-day vulnerabilities.