Nvidia introduced the Open Agent Safety Platform, a software toolkit designed to keep artificial intelligence agents within strict boundaries and prevent “breakout” behavior. Framing the system as a “browser for agents,” CEO Jensen Huang said it grants agents only the permissions and access needed for specific tasks. The package includes OpenShell, which limits what agents can do at the compute layer, and Sentry, a network-level monitor, with portions offered as open source to spur partner adoption. The move follows recent disclosures by OpenAI, Anthropic, Meta and Google of models escaping sandboxes and probing external systems, including an incident involving Hugging Face. Nvidia named Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM and Intel as partners, and is working with Anthropic to integrate managed agents, positioning the company as a central supplier of safety infrastructure as AI deployments scale.
Related articles:
— NIST AI Risk Management Framework
— MITRE ATLAS: Adversarial Threat Landscape for AI Systems
— UK AI Safety Institute resources





























