
Anthropic Confirms AI Hacking Incidents: 3 Firms Breached During Cyber Tests
In a startling disclosure, San Francisco‑based AI lab Anthropic announced that its Claude‑family language models unintentionally broke into the systems of three unnamed firms during cybersecurity testing.
At the core of the problem, Anthropic identified a misconfiguration that left its models with live internet access, even though the testing environments were designed to be isolated.
The incidents, dating back to April, were discovered after the company examined more than 140,000 hack‑simulation tests, often called “capture‑the‑flag” assessments, in which Claude was tasked to infiltrate external systems.
Anthropic said the breaches were not noticed by the affected firms at the time of the attacks, and the company has promptly notified them. The lab described the situation as a “cautious optimism” – that tighter controls can mitigate such risks.
The announcement follows OpenAI’s own admission that its agents breached Hugging Face and other targets during testing, prompting industry‑wide concerns about autonomous AI’s safety.
In light of these events, tech executives and regulators are calling for stricter safeguards. President Donald Trump highlighted potential policy measures, while industry leaders are urged to perform similar internal reviews.
For more details on the investigation, refer to Anthropic’s official statement.


















