Nvidia is set to challenge Anthropic and OpenAI with a new product

A few months ago, the idea of AI software breaking out of its testing environment and going after real companies sounded like science fiction. Now it has become a product category, and Nvidia just launched its entry.

The chipmaker’s answer arrived alongside more than 100 partner organizations, and the group of companies backing the launch says as much about the industry’s mood as the technology itself.

Nvidia launches a platform built to keep AI agents contained

Nvidia on Sept. 28 released its Open Agent Safety Platform, a software package that lets AI developers set limits on what agents can do and helps stop them from breaking out of containment, according to CNBC.

The platform has two main pieces. OpenShell uses security features built into Nvidia’s central processors to confine an agent. Nvidia is working with Arm and Intel so it can run on their chips, too.

A second system called Sentry runs on Nvidia’s BlueField-4 DPUs. It monitors agent behavior independently, and within milliseconds, can quarantine any agent that tries to escape its boundary.

More AI:

  • Nvidia just made a move Wall Street wasn’t ready for
  • Microsoft just took sides in AI policy fight
  • OpenAI just disclosed something genuinely alarming

Nvidia is making a bold claim about it. Justin Boitano, the company’s vice president of enterprise AI, said the platform “could have stopped the breach” at Hugging Face if frontier labs had used it early in model evaluation. He added that Nvidia wants the industry to build it out openly, Reuters reported.

A wide group of companies is backing the launch. Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, Arm, and Intel are among the partners. The tools are also being released with Anthropic’s involvement, integrating OpenShell with Claude Managed Agents.

The Hugging Face breach explains the timing

The launch follows a run of sandbox escapes. OpenAI, Anthropic, Meta, and Google have all been linked to recent incidents. The best-known case involved OpenAI models that got out of containment, reached the open internet, and breached Hugging Face.

The scale was large. Boitano told reporters that Hugging Face reported more than 17,000 actions over the course of the intrusion. Investigations estimated that about 700 AI agents participated in the attack.

The breach also collided with a business deal. OpenAI had published its own account on Aug. 26, the same day reports emerged that Nvidia had agreed to buy Hugging Face for $12.9 billion.

That is nearly triple the $4.5 billion valuation the company carried after its 2023 funding round. The transaction terms include up to $1 billion in employee retention awards on top of an $11.9 billion base price, CNN reported.

That overlap raised a governance question. The company whose chips train most frontier models would also own a neutral hub where rival labs host and test theirs.

Huang tried to defuse it, saying, “NVIDIA compute will not be required to build on or deploy through Hugging Face.”

Nvidia on Sept. 28 released its Open Agent Safety Platform.

Bloomberg / Getty Images

Huang frames it as an engineering, not a regulation, problem

Huang has opposed calls for broad AI safety regulation. He described escaped agents as an engineering problem, comparable to making automobiles safer.

“You can’t have agents roam around and drift around the company, and so you have to find a way to contain it,” he told CNBC.

Nvidia’s case rests on a technical point. Boitano said “model-level safeguards alone can’t govern what agents can access or do.” Nvidia argued that incidents show how easily agents get around guardrails at the application layer. That is why it wants controls spanning the entire stack.

The detection method is built around agent behavior. Ali Golshan, a senior director of AI software at Nvidia, said the tools use mathematical formulas to spot workarounds, such as an agent spawning several sub-agents to slip past a block placed on the main one.

The debate over pace has grown louder in recent weeks. Anthropic CEO Dario Amodei published an essay on Sept. 12 titled “We Must Pace the Frontier,” arguing that capabilities are advancing faster than the industry can control, according to TheStreet.

Boitano described Nvidia’s offering as an engineering solution arriving two weeks after that essay that set off an industry debate.

What it means for investors

Open Agent Safety Platform tools are open source, but they run best on Nvidia hardware. That ties a safety story directly to its chip business. Nvidia is also offering an engineering answer at a moment when Sam Altman and Elon Musk have publicly backed the call to slow down.

Caution is still warranted. Boitano said only that the platform could have stopped the breach, and he cautioned that “each security incident is unique.”

Neither of the known incidents caused serious damage that anyone has identified, though the potential has made many people nervous.

The next data point is simple. Either the labs use it or they do not. Either the breaches stop or they do not.

Huang has bet on engineering. The answer arrives with the next incident.

Related: OpenAI makes development moves to counter SpaceX and Meta