A senior safety researcher at Anthropic warned there is more than a 10% chance advanced AI could wipe out humanity within the next decade, escalating a debate over how quickly to tighten guardrails on frontier systems. Evan Hubinger, who works on AI alignment, said current risks are low but argued the industry lacks a credible plan to control superintelligent models. His comments followed the resignation of Anthropic researcher Jacob Coxon, who claimed leading labs are acting irresponsibly as systems approach superhuman capabilities. The flare-up coincides with rising government scrutiny: the U.K.’s AI Security Institute has reportedly been denied access to Anthropic’s latest model, while a former Treasury minister urged Prime Minister Andy Burnham to pursue a multinational AI safety treaty. Industry figures, including senior leaders at OpenAI and Meta, have recently disclosed instances where autonomous AI agents conducted cyberattacks, reinforcing calls to slow development and strengthen evaluations. Investor interest remains intense amid speculation over potential listings, but experts warned political and regulatory risks are mounting as developers admit they are not “on track” to align superintelligence.
Related article:




























