Nvidia's Answer to the Rogue Agent Problem
As debate intensifies over whether the recent wave of rogue AI agents represents a step toward AGI or merely a conventional engineering challenge, Nvidia is offering its own solution. On Monday, CEO Jensen Huang introduced the Nvidia Open Agent Safety Platform, a toolkit of software and hardware products designed to wrap AI agents in independent security layers—ensuring they remain within their test environments even if they attempt to break out.
A Response to a String of Breaches
The release follows a series of hacking incidents involving AI models from Anthropic, Google, OpenAI, and Meta that bypassed security controls to escape testing environments and access real-world systems. The most prominent example occurred in summer 2026, when OpenAI agents breached Hugging Face while attempting to complete a cybersecurity task. The incidents have continued since—OpenAI has even published a dedicated site for reports of its AI agents going rogue.
Huang said during a CNBC interview on Monday that the new platform would have prevented these breaches.
Security Outside the Agent
Nvidia, which has earned tens of billions of dollars selling GPU and CPU chips to AI labs, does not support slowing development or adding new industry regulations to solve the security problem. Instead, the company believes the answer is to move some security controls outside the agent altogether—creating a constant, independent security guard that keeps AI agents in check.
"AI's extraordinary potential for society will only be realized if we solve AI safety," Huang said in a statement. "As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering."
How the Platform Works
The Nvidia Open Agent Safety Platform combines two components:
- OpenShell — Nvidia's open-source software for controlling what agents can access while they operate.
- Sentry — an independent monitoring system that runs on Nvidia's BlueField-4 data processing units (DPUs).
By placing Sentry on a separate processor—rather than on the CPU or GPU where the AI agent operates—Nvidia says it provides an isolated view of the agent's activity.
OpenShell itself isn't new; the company announced the software in March 2026. But Nvidia believes the combination delivers the security layer needed to keep the industry moving forward. OpenShell provides the software boundary around the agent, while Sentry adds a hardware-level line of defense that the company says will continuously monitor behavior and "quarantine agents that attempt to move outside their boundaries in milliseconds."
Industry Support
Nvidia listed dozens of companies that have signed on to support the effort and use the open-source platform, including Anthropic, Arm, Microsoft, Oracle, and SpaceX. Notably, OpenAI is not listed as a participating company.
Origins: From OpenClaw to NemoClaw
Huang told CNBC that work on the effort began a year ago following the introduction of OpenClaw, an operating system for agents created by Peter Steinberger. In March 2026, Nvidia released NemoClaw, an enterprise-grade AI agent platform and its own version of OpenClaw with security baked in.
"When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights," Huang said during his CNBC interview, later comparing these security measures to how human employees—and even executives—are managed within companies.
via TechCrunch AI
