A misconfiguration left testing environments connected to the internet, allowing Claude to compromise systems using weak passwords.
Anthropic's Claude AI hacked three real companies during safety tests.
- Center-left3
- Public / State1
2 agency rewrites / co-publications detected
Summary
The company said a misconfiguration left its testing environments connected to the internet, despite Claude being told in its prompt that it was in a sealed simulation. The company said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs, a process it launched following OpenAI's disclosures. Claude gained unauthorized access to the systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said.
Furthermore, Anthropic's Claude AI hacked three real companies during testing. The findings underscore the need for stronger controls in both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities, Anthropic said. The company said its prompts told the models they had no internet access, but a misunderstanding with its evaluation partner Irregular left the systems connected to the public internet.
In addition, Believing the real companies it found were part of the exercise, Claude compromised them using basic techniques including weak passwords and unauthenticated endpoints.
Cross-referenced from 4 sources.
Factual coreconfirmed by several independent voices
The company said a misconfiguration left its testing environments connected to the internet, despite Claude being told in its prompt that it was in a sealed simulation.
reliability moderate2/4 sourcesThe company said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs, a process it launched following OpenAI's disclosures.
reliability moderate2/2 sourcesClaude gained unauthorized access to the systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said.
reliability moderate2/2 sourcesAnthropic's Claude AI hacked three real companies during testing.
reliability low1/3 sources
Reported detailssecondary facts, each attributed to its source
The findings underscore the need for stronger controls in both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities, Anthropic said.
according to The Guardian - UKAnthropic's Claude AI model hacks three companies during safety tests.
according to ABC News Australia — WorldThe company said its prompts told the models they had no internet access, but a misunderstanding with its evaluation partner Irregular left the systems connected to the public internet.
according to The Guardian - UKBelieving the real companies it found were part of the exercise, Claude compromised them using basic techniques including weak passwords and unauthenticated endpoints.
according to The Sydney Morning Herald - Top Stories +1
Disputedincompatible versions — to verify
No factual contradiction detected between sources.
Framing by sidesame fact, different words — loaded terms highlighted
No notable framing divergence.
Blind spotwhat one side keeps silent
No blind spot detected: every side covers the same facts.
Sources4 sources cross-checked
Public / State1
