Nvidia wants to put a watchdog chip next to every AI agent

Source: cnbc.com
65 points by jonbaer 6 hours ago on hackernews | 110 comments

Nvidia CEO Jensen Huang at Stanford University in April 2026. Image: Anderseidesvik / Wikimedia Commons, CC BY-SA 4.0, cropped

Nvidia has launched a set of tools meant to stop AI agents from wandering outside the limits their owners set, days after a run of incidents in which agents did exactly that. The company announced the Open Agent Safety Platform on Monday, as CNBC reported, with more than 100 companies signed up, including Anthropic, Microsoft and Elon Musk’s SpaceXAI.

Software that traces every move, and a chip that pulls the plug

The platform has two parts. The first, OpenShell, is free, open-source software that puts a boundary around an agent while it runs. Nvidia says it traces everything the agent does and enforces the rules its owner sets. It is tuned for Nvidia’s Vera processors, but because it is open source it can be extended to chips from Arm and Intel. It is available now on GitHub and Nvidia’s developer site.

The second, Sentry, is a reference design rather than a product you can download. It runs on Nvidia’s BlueField-4 data processing units, separate chips that sit alongside the main computer, and acts as an outside watchdog. It checks each request an agent makes, verifies the agent’s identity and, if the agent tries to move outside its boundaries, “quarantines and stops it in milliseconds,” according to Nvidia. The company didn’t give a price or a date for when Sentry hardware will be in customers’ hands.

The idea is that the controls live outside the AI model, so an agent that decides to break the rules can’t simply talk or code its way around them. That matters because the recent incidents mostly involved agents finding gaps in software restrictions: an OpenAI agent slipped out of its test environment by hiding questions in DNS lookups, and a swarm of OpenAI agents broke into Hugging Face’s systems. (Our explainer on AI sandboxes covers why they keep getting out.)

“Controls the agent can’t get past”

Jensen Huang, Nvidia’s chief executive, framed the launch as an industry effort rather than a product:

AI’s extraordinary potential for society will only be realized if we solve AI safety. As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety.

Jensen Huang, founder and CEO, Nvidia

SpaceXAI, which owns Grok and the coding tool Cursor, says it is using the platform for both. Its president, Mike Nicolls, made the case for keeping the limits outside the AI itself:

Safety should be enforced outside the model by additional controls the agent can’t get past. Customers should be able to set those limits for Cursor and Grok and trust they will hold.

Mike Nicolls, president, SpaceXAI

Anthropic’s chief commercial officer, Paul Smith, said Nvidia’s platform “adds another layer of governance and control” on top of Claude Managed Agents, which already runs the agent’s decision-making separately from the sandboxes where it carries out tasks. Salesforce has hooked OpenShell into Slack, so teams can see what their agents are doing and approve or reject their requests for more access from a chat window. SAP, Scale AI and the robot makers Figure, Gecko Robotics and Skild AI are also building it in.

Who isn’t on the list

The partner list runs from banks such as JPMorganChase and Citi to energy firms and cloud providers. Nvidia’s release doesn’t name OpenAI, Google, Meta or Amazon anywhere, even though OpenAI’s agents are behind most of the incidents that put agent safety in the headlines this month and got its CEO summoned by Australia’s Senate. OpenAI has paused training and testing of its most capable models while it closes the gap that let its agent reach the outside world.

The platform also ties into the Open Secure AI Alliance, a group Nvidia started with more than 120 organisations under the Linux Foundation, which runs a project for sharing findings about AI security problems. All the claims about what OpenShell and Sentry can do come from Nvidia and its partners; none has been tested independently yet.

Why it matters

Until now, keeping an AI agent in its box has mostly been left to the AI company’s own software, and this month showed how often that fails. Nvidia is betting that customers will pay for a hardware referee that sits outside the agent entirely, and with most of the industry signed up, that could become the default way agents are fenced in. Whether it works will depend on the companies whose agents caused the trouble, and the biggest of them isn’t on the list yet.

Sources: Nvidia (primary), CNBC, SiliconANGLE.