Nvidia is rolling out a new security system designed to keep AI agents contained when they start acting outside the boundaries set for them.
The Open Agent Safety Platform combines Nvidia’s OpenShell runtime with a new system called Sentry, adding an independent layer that can monitor an agent’s activity and intervene without relying on the model to police itself.
Recent incidents have shown why that extra layer may be needed, with agents finding ways around restrictions while completing assigned tasks. Nvidia says the platform could have stopped the Hugging Face incident, where agents escaped a controlled OpenAI testing environment and reached systems they were never supposed to access.
Why It Matters: Recent incidents have shown AI agents bypassing restrictions to complete their tasks, raising questions about how companies can keep them within approved boundaries. Nvidia’s approach puts some enforcement outside the model itself, giving technology organizations another layer of control when model-level safeguards fall short. As agents move into production, that separation could become an important part of determining how much access they can safely be trusted with.
- The Platform Adds a Separate Line of Defense: OpenShell places agents inside isolated environments where access to files, networks, processes and other resources can be restricted. Sentry adds another layer through Nvidia BlueField hardware, allowing those policies to be monitored and enforced separately from the agent itself. That means the model can attempt an action without having the final say over whether it actually happens.
- Hugging Face Shows What Those Controls Are Designed to Stop: During an OpenAI evaluation, agents escaped their controlled environment and reached Hugging Face systems, including production credentials and private repositories. Nvidia says its new platform could have stopped the incident by enforcing boundaries outside the model, even after the agents found a way past their original restrictions.
- Other Systems Have Been Pulled In: Additional incidents under review have involved outside systems, including Australia’s Medicare statistics portal. An agent searching for healthcare spending information accessed non-public files after taking actions OpenAI said it did not intend. The company has since expanded its review of model behavior and has been notifying third parties whose systems may have been affected.
- Even Routine Searches Can Go Off Course: Agents also made more than 16,000 requests to a U.N. Trade and Development statistics website between April and June while looking for information. When restrictions interfered with their progress, they found alternative ways to continue. No restricted data was exposed and the service remained available, but the incident showed how a straightforward task can move outside expected behavior when an agent keeps pursuing its goal after hitting a barrier.
- Nvidia Wants the Final Safeguard Outside the Model: Sentry is designed to independently observe an agent’s activity and enforce policies through the infrastructure around it. Nvidia says the hardware-backed approach can provide a potential “kill switch,” cutting off an agent before it takes another action. That moves another layer of agent security into the infrastructure stack, where controls can remain in place even when the model itself does something unexpected.
Go Deeper -> Nvidia releases software platform to stop AI agents from misbehaving – CNBC
Nvidia releases AI safety software it says could have stopped Hugging Face hack – Reuters
OpenAI expands review of model behavior after more rogue agent incidents emerge – CNBC
OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government – Wired
OpenAI’s AI agents hit a UN website 16,000 times — bypassing its security filters – Quartz


