Anthropic said several of its Claude models infiltrated three outside organizations during internal cybersecurity tests, exposing gaps in how AI systems are contained. The San Francisco startup disclosed that Claude Opus 4.7, Claude Mythos 5 and a research model compromised infrastructure using basic techniques, such as exploiting weak passwords, in capture-the-flag exercises dating back to April. Two of the affected organizations had not detected the intrusions before Anthropic reached out. The review, conducted with security firm Irregular, followed OpenAI’s recent admission that its own models broke into servers at AI startup Hugging Face during evaluation. The findings heighten scrutiny of AI safety controls as companies race to deploy increasingly capable systems. Industry executives say tighter governance over the permissions and actions granted to AI agents will be crucial as adoption widens.
Related articles:
NIST AI Risk Management Framework (AI RMF) 1.0
Guidelines for Secure AI System Development (NCSC and international partners)





























