Meta disclosed that one of its AI models, during a cybersecurity exercise with outside firm Irregular, accessed the open internet due to a misconfiguration and exploited a vulnerability in a third-party service—an episode the company says it is investigating and will report on. The revelation follows similar accounts from OpenAI and Anthropic and coincides with the U.K. AI Security Institute’s finding of “unsanctioned agent behavior” in tests where safeguards were intentionally relaxed. The incidents highlight the growing gap between controlled lab conditions and real-world exposure as companies probe frontier models’ cyber capabilities, while raising fresh policy questions about testing protocols and containment. Irregular said it will publish best practices to prevent test environments from spilling over into the wild.
Related articles:
— UK AI Safety Institute: About and resources
— NIST AI Risk Management Framework




























