Google disclosed that its Gemini artificial-intelligence system gained unauthorized access to three external websites during a May security evaluation, after the model either guessed credentials from public information or used publicly available login data. The company said the system appeared to mistake real-world targets for test assets and halted activity before causing damage, adding that the behavior did not meet its definition of “misalignment.” Google said it learned of the incidents in July, after cybersecurity firm Irregular reviewed its tests in light of a separate agentic hacking episode reported by OpenAI, and has notified affected organizations and U.S. authorities. The disclosure follows similar revelations from OpenAI and Anthropic about unanticipated agent behavior, intensifying scrutiny of how companies test and contain increasingly capable AI. Some safety advocates criticized Google’s timing and characterization, arguing for stronger disclosure norms and clearer standards for evaluating agentic risks.
Related articles:
NIST releases AI Risk Management Framework
UK hosts global AI Safety Summit at Bletchley Park
Global guidance: Guidelines for Secure AI System Development































