Anthropic disclosed that its Claude artificial-intelligence models gained unauthorized access to three companies’ systems during internal cybersecurity evaluations after a configuration error allowed internet access in test environments meant to be isolated. The company said it reviewed 141,006 test sessions beginning July 23 and identified three incidents by July 24, notifying the affected organizations on July 27. Two of the firms were unaware of the activity until Anthropic reached out. The breaches occurred during capture-the-flag exercises and involved three models—Claude Opus 4.7, Claude Mythos 5 and an internal research model—that exploited basic security lapses such as weak passwords and unauthenticated endpoints. Anthropic attributed the exposure to a misunderstanding with its evaluation partner, Irregular, which left systems connected to the public internet despite prompts indicating no connectivity. The admission follows OpenAI’s recent report that an autonomous agent compromised infrastructure at AI firm Hugging Face during a separate test, underscoring rising concerns that increasingly capable AI systems can execute real-world cyber actions. Anthropic said the findings highlight the need for tighter controls and safeguards in both internal and third-party testing environments as model capabilities advance.
Related articles:
— Guidelines for Secure AI System Development (NCSC/CISA)
— NIST AI Risk Management Framework
— OWASP Top 10 for Large Language Model Applications
— MITRE ATLAS: Adversarial Threat Landscape for AI Systems





























