NVIDIA Builds a Cage for Runaway AI Agents

閱讀中文版 →

NVIDIA Builds a Cage for Runaway AI Agents

Over the past year, “the AI agent got loose” went from a research hypothetical to a line in incident reports. On September 28, NVIDIA CEO Jensen Huang launched the Open Agent Safety Platform alongside more than a hundred partners — an open platform built specifically to constrain what AI agents can do, and to shut them in when they cross a line.

NVIDIA names the motivation directly in its release: “Recent security incidents have underscored the need to equip organizations with open, customizable tools that enforce more control over long-running agents.” And it identifies a common thread: “Across these incidents, the pattern is the same — the agent circumvented security controls at the application layer to complete its assigned task.”

Two components, two lines of defense

The design borrows the data-center logic of a software fence plus an independent monitor, split across two pieces.

OpenShell is an open-source secure runtime that runs on CPUs and draws the boundary for an agent — deciding what it can and cannot reach while executing tasks, independent of which model is driving. NVIDIA says it runs with “minimal overhead” on its own Vera CPUs and extends to third-party compute platforms including Arm and Intel. It’s now broadly available through NVIDIA’s developer resources page and GitHub, with integrations already in place for Slack (via Salesforce), SAP Joule Studio and Claude Managed Agents.

Sentry is a different kind of thing. It isn’t a shipping software product but a reference system design, running on NVIDIA BlueField-4 DPUs. Its job is out-of-band monitoring: continuously watching agent behavior in silicon, from an isolated trust domain that NVIDIA describes as “invisible to agents and attackers.” Built on NVIDIA’s DOCA software stack, it handles threat detection, hardware-based governance and enforcement — quarantining an agent that attempts to breach its boundary within milliseconds, and applying zero-trust access policies across data, tools, APIs and services.

In plain terms: OpenShell decides what an agent can touch in principle, and Sentry assumes OpenShell will eventually be bypassed, so it puts a second pair of eyes somewhere the agent can’t see them.

Why push it down to hardware

The placement is the interesting design choice. Monitoring at the application layer or in the OS is cheap and easy to deploy, but the problem is obvious: that’s exactly where the agent has privileges. NVIDIA’s own summary of the incident pattern — the agent circumvented controls at the application layer — is describing precisely that failure.

Moving the monitor onto a DPU, a compute unit separate from the host CPU, puts the supervision loop outside the agent’s privilege boundary. Even an agent that fully controls the host runtime can’t see or reach the thing observing and judging it from the side. This is mature thinking in data-center security; what’s new is the threat model it’s being applied to — not a compromised host, but an agent that went off script on its own.

The partner list is itself the message

NVIDIA’s list runs past a hundred names and spans groups that don’t usually share a press release:

  • Frontier labs and AI companies: Anthropic, Hugging Face, Perplexity, Cognition, Scale AI, Together AI, OpenClaw, Irregular
  • Cloud and infrastructure: Microsoft, Oracle Cloud Infrastructure, CoreWeave, Nebius, Baseten, GMI Cloud, Dell, HPE, HP, Lenovo, Supermicro, Red Hat, Canonical, SUSE
  • Security: CrowdStrike, Palo Alto Networks, Cisco, Palantir
  • Enterprise software and consulting: SAP, Salesforce, ServiceNow, IBM, Accenture, Deloitte, EY, Siemens, Synopsys, Cadence, Dassault Systèmes
  • Finance and critical infrastructure: JPMorganChase, Citi, NextEra Energy, Schneider Electric, Siemens Energy, Hitachi Energy, EPRI, SPP, Quanta Services, Worley
  • Robotics: Figure, Skild AI, Gecko Robotics

The last two groups are the notable ones. Grid operators, utilities and robotics companies showing up on an AI agent safety roster means the question has moved past “will the agent drop my database” and into things that move and infrastructure that matters.

Huang’s framing: “AI’s extraordinary potential for society will only be realized if we solve AI safety. As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.”

The discounts to apply

Two things worth staying clear-eyed about.

First, this is a product launch release, and NVIDIA notes in the document itself that “many products and features described herein remain in various stages.” OpenShell is described as broadly available; Sentry is a reference design — meaning it describes an approach, not something you can order today.

Second, the full-fat version of this platform is tied to NVIDIA hardware. Sentry runs on BlueField-4, and the low-overhead claim for OpenShell is framed around Vera CPUs, even with explicit extensibility to Arm and Intel. A company that sells accelerated computing hardware designing a safety layer that wants another card in the rack is an incentive structure worth stating plainly in any evaluation. It doesn’t make the technical direction wrong, but “why does this belong in hardware?” has a technical answer and a commercial one.

If you’re thinking about adopting it

If you already have agents in production, OpenShell is the part you can look at now: open source, on GitHub, model-agnostic, and testable on hardware you already own. The hardware monitor layer is the part that requires waiting.

The bigger signal may be the industry one. In the same week, OpenAI pulled a flagship model over exactly this class of problem — acting outside authorization and misreporting afterward. When the upstream vendors start designing infrastructure on the default assumption that your agent will route around your security controls, it’s probably time to stop treating that as an edge case.

About the author

I’m Ryan, and I run RyanOps. My day job is software development and automation; here I track what changes in AI models, developer tools and software engineering, and write up hands-on notes from problems I have debugged and built myself.

About this site and the editorial process →