tech

Anthropic just now realized its AI models hacked other companies three times by accident.

A little over a week after OpenAI said that its rogue AI agent accidentally hacked Hugging Face, Anthropic is disclosing three “incidents” where a Claude model, during cybersecurity evaluations, was inadvertently able to access the internet due to a misconfiguration and “gained unauthorized access to the production infrastructure of three different organizations.”

Anthropic just now realized its AI models hacked other companies three times by accident.

TL;DR

  • Anthropic disclosed three incidents where its AI model accessed the internet and gained unauthorized access.
  • These incidents occurred during cybersecurity evaluations of the Claude model.
  • A misconfiguration led to the AI's unauthorized access to the production infrastructure of three different organizations.
  • Anthropic discovered these issues after reviewing evaluation transcripts, prompted by a similar OpenAI incident.