A U.K.-based nonprofit, Tech Against Terrorism, found that three in five AI models failed a terrorism safety test designed to probe whether systems would assist in planning mass-casualty attacks. The group evaluated more than 130 models and deemed a system to have failed if it provided a specific, actionable response once or scored below 90 out of 100 on its safety benchmark.
Open-weight models performed comparably to closed systems in baseline tests, but every model subjected to “abliteration”—the removal of safety guardrails by modifying weights—failed, sometimes catastrophically. An abliterated version of Meta’s Llama 3.1 8B, which initially scored 97, dropped to roughly 3 and produced detailed responses to violent prompts, the report said. Similar behavior appeared in abliterated builds of other popular open-weight models. Tools to strip safeguards are free and widely available, and repositories advertised thousands of “uncensored” models, according to the group.
Meta and Hugging Face said they enforce policies and moderation; Hugging Face cautioned that blocking potentially harmful requests can also impede legitimate research. The nonprofit, backed by several governments and supported by the U.N. Counter-Terrorism Directorate, urged funding for independent benchmarks, hardening models against abliteration, and restricting distribution of stripped versions—while arguing safety and open development can coexist.
Related articles:
Meta Llama 3
Universal and Transferable Adversarial Attacks on Aligned Language Models
NIST AI Risk Management Framework































