
Nvidia has announced a policing system that will watch and rein in artificial intelligence (AI) agents before they go rogue. The Open Agent Safety Platform combines open source software and a reference system design to keep AI agents within set boundaries from testing through deployment.
Recent high-profile incidents have shown that agents are creative in finding unintended ways to achieve the goals they've been given and cannot be expected to police their own behavior, said Justin Boitano, vice president and general manager of enterprise computing at Nvidia, at a press briefing on Monday. Frontier AI labs have reported that agents have circumvented security controls to escape the evaluation environments meant to contain them, accessed systems they are not authorized to use, and, in some cases, failed to report their actions. Agent drift — when an agent deviated from its intended task — occurs when an agent encounters a policy block, a software bug, or a missing tool. Drift can also happen if the agent receives ambiguous instructions or has been running for days or weeks trying to solve a hard problem.
"Once AI can act, safeguards must govern the agent's actions," Boitano said.
Nvidia's Open Agent Safety Platform integrates two security layers that sit between enterprise IT systems and AI agents to define, monitor, and contain agents. One layer is OpenShell, an open source runtime that Nvidia originally introduced in March to sandbox agents and enforce policy on what they do. OpenShell supports Codex, Claudo Code, Pi, Hermes, and other agents and runs on some x86 and Arm CPUs. The second is Sentry, a hardware-based watchdog layer that runs on Nvidia's BlueField-4 data processing units (DPUs) and monitors agents' actions.
"With OpenShell, the security team can formally verify an agent has enough authority to do its job and no more," Boitano said. "If a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly."
OpenShell Enforces Agent Controls
OpenShell has three components: a gateway that manages sandbox life cycles and policies, a sandbox that applies kernel-level controls to file system and process activity, and a supervisor for each sandbox that checks outbound requests against policy. For example, the supervisor inspecting network traffic can allow an agent to read data through an API but block all write attempts. The supervisor can enforce controls even when the agent is running code it generated during a task, and all policy decisions are logged, according to Nvidia.
Agents can propose policy changes if the policy adviser feature is enabled, but agents cannot approve their own requests, ensuring human oversight. A formal logic policy prover also checks whether the permissions a policy grants, as modeled, stay within limits set by the operator.
Sentry Provides Hardware Enforcement
Hardware enforces the agentic policing system. Nvidia's Arm-based Vera CPU runs OpenShell, which sets boundaries in which agents can run safely. The BlueField-4 data processing unit, which also includes an Arm CPU, provides the horsepower Sentry needs to observe agentic behavior and protect systems.
Sentry provides in-silicon security enforcement by observing agent activity and blocking attempts to move outside the software boundary, Nvidia said. Because the data processing unit operates separately from the agent's host, Sentry can still block agents even when the host is compromised. Sentry is built on Nvidia's DOCA software, which inspects agent requests and responses, provides attested telemetry, verifies agent identities, and enforces zero-trust access policies for data, tools, APIs, and services.
Not Trusting the AI Agent to Stay in Bounds
These types of limits are necessary because Nvidia's own tests found that frontier agents with reduced safeguards could spend up to two hours trying to convince an AI reviewer to grant permissions to modify a protected GitHub repository.
"The organization should not have to trust the agent to respect that boundary," Boitano said. "The infrastructure should enforce it explicitly."
During the briefing, Boitano said that using an AI agent policing system during model testing could have prevented the July incident when OpenAI agents infiltrated Hugging Face servers.
Agents going rogue and breaking into systems is an "eventuality," and enterprises need a robust system to inspect and audit AI agents and systems, says Jack Gold, principal analyst at J. Gold Associates.
"That requires an ability to monitor and control actions not only at a local level, but also ones that interact across a network of compute, data, and application resources," he says.
Nvidia is working with Anthropic to implement the Open Agent Safety Platform, while SpaceXAI is already using the technology. Salesforce and Nvidia integrated OpenShell with Slack to allow teams to view agent activity, audit events, and approve or reject agent requests for more permissions. SAP is planning to embed OpenShell in its JouleStudio runtime. Other partners include CrowdStrike, Microsoft, Perplexity, Accenture, ServiceNow, and IBM to bring the technology to customers.
Organizations already running Vera systems with BlueField-4 can enable Sentry's protection layer with a software update. OpenShell and related skills are available via Nvidia’s developer resources page and on GitHub.
"The industry does not need agents that promise to stay within bounds," Boitano said. "It needs systems that can prove and enforce those boundaries."