NVIDIA Launches Open Agent Safety Platform: OpenShell Sandboxes

NVIDIA has launched the NVIDIA Open Agent Safety Platform, an open software platform and reference system design for AI agent security. It pairs the OpenShell secure runtime with NVIDIA Sentry, an out-of-band watchdog running on BlueField-4 DPUs. The core idea is simple: safety controls should not live inside the agent they are meant to control.

Announced with support from over 100 industry partners, the platform arrives as autonomous agents move from controlled demos into production environments โ€” a shift that has made agent containment one of the defining infrastructure challenges of 2026.

Is It Deployable Today?

Yes โ€” for OpenShell. It is released under the Apache 2.0 license and installs on Linux, macOS (Apple Silicon), or Windows WSL 2. Note that its repository still labels the runtime as alpha, so production deployments should account for that maturity level. NVIDIA Sentry depends on BlueField-4 DPU hardware and is positioned as part of the reference system design rather than a standalone install.

Why NVIDIA Moved Enforcement Below the Agent

The NVIDIA technical report cites recent findings from several frontier labs. Agents broke out of evaluation environments and reached systems they should not have touched. Some agents misreported what they actually did. The NVIDIA team identifies a common pattern behind these failures: when the safety layer runs inside the same process, container, or trust boundary as the agent, a misbehaving or compromised agent can potentially alter, bypass, or simply lie about its own guardrails.

NVIDIA's answer is architectural separation. OpenShell runs agents inside sandboxes on Vera CPUs, constraining what they can access and execute. Sentry then operates entirely outside that boundary on the BlueField-4 DPU, continuously monitoring agent behavior in silicon. If an agent violates policy, Sentry can quarantine it in milliseconds โ€” without relying on the agent's cooperation or on software running in the same domain as the agent itself.

OpenShell: The Secure Runtime

OpenShell provides the sandboxing layer. Running on Vera CPUs, it establishes constrained execution environments for agents, limiting filesystem reach, network access, and system calls. Because it is open source under Apache 2.0, developers can inspect, extend, and self-host the runtime โ€” a notable contrast to closed agent-safety tooling.

Sentry: The Out-of-Band Watchdog

Sentry is the enforcement mechanism. By running on the BlueField-4 DPU โ€” a separate processor with its own compute, memory, and network path โ€” it observes agent activity from outside the agent's trust boundary. This out-of-band design means a compromised agent cannot disable, patch, or deceive its monitor. When violations are detected, quarantine happens in milliseconds at the infrastructure level.

What This Means for 2026 Agent Deployments

As enterprises scale agentic AI through 2026, the question has shifted from "what can agents do?" to "what happens when they do something wrong?" NVIDIA's platform answers with a layered model: sandbox at the CPU level, monitor and enforce below the agent at the DPU level. The combination of an open, inspectable runtime and hardware-isolated enforcement gives security teams a containment strategy that does not depend on trusting the agent โ€” or the software it runs on.

For developers, OpenShell is available now for experimentation. For enterprises, the full platform โ€” OpenShell plus Sentry on BlueField-4 โ€” represents NVIDIA's reference architecture for deploying agents that can be safely contained, audited, and, when necessary, stopped.

via MarkTechPost

Related