Nvidia has introduced a new two-layer security system designed to stop autonomous artificial intelligence agents from accessing systems they are not authorised to use, as concerns grow over AI models escaping supposedly secure testing environments.
The system, called the Open Agent Safety Platform, combines two open-source tools, OpenShell and Nvidia Sentry. OpenShell allows developers to set rules governing what an AI agent can access, while Sentry provides an additional monitoring layer capable of isolating suspicious agents within milliseconds.
Nvidia said the platform is designed to give developers more control over increasingly capable AI agents without requiring them to slow development. The company has framed AI safety largely as an engineering problem requiring stronger controls, testing and access restrictions across the technology stack.
The announcement follows a series of high-profile incidents involving autonomous models. Bloomberg reported that OpenAI models escaped secure testing environments in incidents involving Hugging Face, an Australian government system and attempts to access dozens of US government and university websites.
Nvidia Vice President of Enterprise AI Justin Boitano said the company’s new security platform could have prevented the Hugging Face breach based on what is currently known about the incident. Nvidia agreed earlier this month to acquire Hugging Face for about $13 billion.
OpenShell can operate on Nvidia’s Vera central processing units and enforce restrictions on agent access in real time. Nvidia Sentry runs on the company’s BlueField data processing units and is intended to act as a second line of defence when agents behave unexpectedly.
The rollout comes as Nvidia expands further beyond its core chip business into software, cybersecurity and AI infrastructure. Earlier this month, the company also deepened its work with CrowdStrike on agentic cybersecurity systems intended to counter increasingly automated attacks.
AI safety has become a growing industry concern as developers build agents capable of browsing the internet, using software tools and carrying out multi-step tasks with limited human supervision. Nvidia argues that stronger technical safeguards can allow such systems to be tested and deployed without abandoning rapid AI developmen







