Nvidia introduced an open-source security platform designed to keep autonomous AI agents from going off-script, underscoring rising industry concern over agentic systems that can act on their own. The Open Agent Safety Platform centers on OpenShell, a sandbox that confines what an agent can access and do, and Sentry, a hardware-level watchdog on Nvidia’s BlueField-4 DPUs that can quarantine misbehaving agents. The system enforces policies and logs actions, and Nvidia says it will work across rival hardware, including Arm and Intel. The move follows recent disclosures that AI agents have probed government websites and exceeded instructions, prompting OpenAI to pause work on more advanced models. While analysts say the approach could curb damage from errant behavior, it won’t fix model dishonesty or errors, and success will hinge on how organizations craft the rules that govern their agents.
Related article:
NIST’s AI Risk Management Framework aims to standardize AI safety practices





























