AI agent security

NVIDIA Open Agent Safety: why agent controls are moving outside the model

•Make Better Editorial

NVIDIA’s Open Agent Safety Platform combines OpenShell runtime policies with an independent Sentry monitoring layer. The key design lesson: production agents need enforceable operational boundaries.

NVIDIA has launched an Open Agent Safety Platform built around a simple operational assumption: autonomous agents should run inside boundaries enforced by systems outside the agent itself. The platform combines OpenShell, an open-source runtime that isolates agents and applies policy, with Sentry, an optional independent monitoring layer designed to observe and stop activity outside those boundaries.

What NVIDIA launched

Two control layers

OpenShellSentry
Secure runtime around the agentIndependent monitoring layer
Isolated execution and policy enforcementRuns on NVIDIA BlueField-4 DPUs
Controls access to files, networks, tools and credentialsMonitors behavior and can stop activity outside defined boundaries
Open source and designed to extend beyond NVIDIA-only computeHardware-backed reference design for organizations needing another trust layer

NVIDIA says OpenShell turns operator requirements into runtime policy and applies those limits while the agent works. Its technical documentation describes isolation, monitoring and policy controls around resources the agent can reach. Sentry adds a separate enforcement domain that the agent does not control, using BlueField hardware to monitor activity independently.

The important design shift: policy outside the model

Make Better analysis

Instructions about what an agent should or should not do are useful, but they are not the same as technical permissions. A production agent can encounter ambiguous instructions, untrusted content, bugs or tool failures. The stronger architecture is to let the model decide how to complete its task only inside deterministic limits enforced by another layer.

Practical principle

Separate the autonomous worker from the system that decides what resources and actions are permitted. High-value permissions should remain enforceable outside the agent loop.

What should be enforced at runtime

Instruction vs enforceable control

Agent instructionRuntime control
Only use approved dataRestrict filesystem, database and API scopes
Keep sensitive information inside approved systemsRestrict network destinations
Protect credentialsKeep secrets outside the agent workspace and provide scoped access
Use only approved toolsLimit processes and execution capabilities
Stop when behavior leaves the expected scopeMonitor activity and terminate the run externally

You do not need specialized hardware to adopt the core pattern

Sentry is the hardware-backed layer in NVIDIA's reference architecture, but the broader lesson does not depend on buying a particular system. Teams can already separate an agent from its permissions using isolation, scoped service accounts, network policies, secrets management, approval gates and external audit logs. OpenShell itself is open source, and NVIDIA says it can be extended to third-party compute platforms.

A minimum control stack for production agents

  1. Run the agent in an isolated environment rather than directly on a highly privileged production host.
  2. Start with minimal access and explicitly grant only the files, APIs, networks and tools required for the job.
  3. Keep credentials outside the agent's editable workspace and provide narrowly scoped access.
  4. Record tool use and permission decisions in a separate audit system.
  5. Require approval or a separate execution service for high-impact external actions.
  6. Maintain an independent way to stop or isolate a run without depending on the agent itself.

What the announcement does not establish

NVIDIA describes the platform as a stronger architecture for agent containment, but the announcement is not evidence that every deployment using OpenShell or Sentry is secure. Policy quality, integration choices, credentials, application vulnerabilities and the surrounding infrastructure still matter. Claims about how the architecture might have changed earlier incidents should be treated as vendor assessments rather than demonstrated counterfactuals.

Bottom line

The useful idea behind NVIDIA Open Agent Safety is architectural: autonomy and authority should be separated. Let agents reason flexibly inside a constrained environment, while deterministic systems outside the model control what resources they can reach, record what they do and retain the ability to stop them.

Sources & useful resources