NVIDIA’s Open Agent Safety Platform Moves Guardrails Outside the Model

NVIDIA’s OpenShell runtime restricts AI-agent files, networks, processes and credentials, while Sentry adds a BlueField-4 hardware watchdog. Here is how the stack works and where its claims still need proof.
NVIDIA Open Agent Safety Platform graphic featuring OpenShell
NVIDIA’s Open Agent Safety Platform combines the OpenShell runtime with the Sentry hardware reference design. Image: NVIDIA

NVIDIA has launched an agent-security stack that moves enforcement outside the AI model and, in its most complete form, outside the host system running the agent.

Announced on September 28, the NVIDIA Open Agent Safety Platform combines OpenShell, an Apache 2.0 secure runtime that is broadly available now, with a Sentry reference design built around NVIDIA BlueField-4 data processing units. NVIDIA says Sentry can quarantine an agent in milliseconds when it crosses a defined boundary.

The launch addresses a concrete problem with autonomous agents: a model can ignore a prompt-level restriction, generated code can spawn new processes, and a compromised tool can try to reach files or services the operator never intended to expose. OpenShell treats those risks as an infrastructure problem. It limits files, networks, processes and credentials with controls that the model cannot rewrite from inside its own workload.

How OpenShell contains an AI agent

OpenShell runs each agent in a separate sandbox. A local supervisor launches the agent as a restricted child process and evaluates its outbound requests against a declarative policy. The control plane, called the Gateway, stores state, distributes policy and settings, and manages model-provider and credential configuration.

The design matters because agent-generated code inherits the same boundary. According to NVIDIA’s OpenShell documentation, filesystem rules use Linux Landlock restrictions, while unprivileged identities and seccomp filters block privilege-escalation paths and dangerous system calls. Filesystem and process limits are locked when a sandbox is created. Network policy can be updated while the workload is running.

OpenShell also separates an agent’s need for a service from the underlying secret. The workload receives an opaque credential reference rather than a raw API key. The supervisor releases the real credential only after the destination and authorization checks pass. For configured HTTP, GraphQL and Model Context Protocol traffic, policy can distinguish between operations, allowing a read while denying a write through the same API.

That is more specific than telling an agent to “only use approved tools.” A security team can define which directories are readable, which paths are writable, which hosts can receive traffic, which model providers are permitted and which credentials can be used at each endpoint. OpenShell denies undeclared access by default.

Sentry adds a second control point in hardware

The Sentry portion of NVIDIA’s reference design is an optional, out-of-band monitor running on a BlueField-4 DPU. In NVIDIA’s Vera Rubin POD architecture, the DPU sits on the node’s path to the model. That position lets it observe agent interactions and interrupt model access without depending on the host operating system or agent process.

Sentry uses NVIDIA’s DOCA software to correlate agent identity, policy decisions, tool calls and data access. The company describes it as a backstop for cases in which the host is compromised or a workload finds a way around a software control. NVIDIA’s technical overview says the hardware layer can enforce policy at line speed and maintain a contextual record across fleets of agents, subagents and tools.

The two components should not be treated as equally deployable today. OpenShell is open source and can run with open or closed models on supported local, cloud, hybrid, on-premises and air-gapped systems. NVIDIA says it can also be extended to Arm and Intel compute. Sentry is a reference design tied to BlueField-4 for the independent hardware enforcement NVIDIA is promoting.

Who is working with the platform

NVIDIA named more than 100 organizations working with technologies in the platform, a group spanning model developers, security vendors, banks, infrastructure providers and robotics companies. The list includes Anthropic, Microsoft, Cisco, CrowdStrike, Hugging Face, JPMorganChase, Palo Alto Networks, Red Hat, Salesforce, SAP and ServiceNow.

The announcement uses different language for different partners, so the list is not evidence that every company has deployed the full stack. NVIDIA says SpaceXAI is using it with Cursor coding agents and Grok models. Anthropic has collaborated on integrating its Managed Agents architecture with OpenShell and BlueField controls, while Red Hat runs OpenShell and DOCA in its AI Factory offering. Other organizations are described as integrating, supporting or building with parts of the platform.

What the platform cannot solve

OpenShell can enforce a policy, but it cannot decide whether that policy is wise. It does not make a model truthful, correct or aligned with a user’s intent. A rule that is too permissive can still expose sensitive data. A rule that is too narrow can stop an agent from completing legitimate work.

That tradeoff is the central deployment question. University of Wisconsin computer science professor Somesh Jha told The Associated Press that the software could block useful agent behavior and that case studies will be needed to show how well NVIDIA balances capability with security.

There is also a difference between containment and behavioral detection. Filesystem, network and process policies are deterministic: an operation is either allowed or denied. Detecting “drift” from an agent’s intended task requires a behavioral profile and can produce harder judgment calls. NVIDIA has not yet published independent benchmark results showing false-positive rates, escape resistance across varied workloads or the operational overhead of the complete OpenShell-plus-Sentry design.

What teams should test before deployment

Organizations evaluating OpenShell should start with one bounded workflow rather than a general-purpose agent. A useful pilot would give a coding agent access to a single repository, an approved package registry, one model endpoint and temporary credentials that cannot reach production.

  • Write separate read and write rules for repositories, issue trackers and MCP tools.
  • Allow network access only to named endpoints, then test redirects, child processes and generated install scripts.
  • Keep raw cloud, source-control and model-provider secrets outside the sandbox.
  • Record policy changes in version control and require review for new hosts, writable paths or credential attachments.
  • Exercise denial and quarantine paths before production, including how work is recovered and who can restore access.
  • Measure blocked legitimate actions, latency and operator overrides instead of judging the system only by whether an escape test fails.

NVIDIA’s launch makes agent safety look more like conventional zero-trust engineering: isolate the workload, minimize authority, broker credentials and keep an independent enforcement point. OpenShell provides a deployable version of that approach today. The stronger claim, that a hardware watchdog can reliably contain large fleets of long-running agents without making them unusable, still needs evidence from real deployments.

Previous Post
An iPhone on a table, representing Apple security updates and mobile device patching

Apple’s iOS 26 Zero-Day Fix: Update to 26.7.1 Now

Related Posts