Nvidia unveils new system to stop AI agents from misbehaving

28 Sep 2026

Jensen Huang, co-founder and CEO of Nvidia. Image: Fortune Global Forum/Flickr (CC BY-NC-ND 2.0)

Jensen Huang previously stated that he wants the industry to self regulate on AI safety, and that new laws would be unnecessary.

Nvidia has unveiled an open-source software platform that promises to better restrict AI agents from taking problematic actions in order to achieve its goals.

The ‘Open Agent Safety Platform’ aims to strengthen AI security from agent testing to deployment, with software that monitors and sets boundaries for agents running on CPUs.

One part of the stack, the OpenShell software, runs agents in a sandbox and turns operators’ instructions into a defined policy. The software is compatible with third-party compute platforms, including those from Arm and Intel.

Meanwhile, Sentry provides in-silicon security enforcement, giving it the ability to quarantine and stop misbehaving AI agents.

The launch is a direct response from Nvidia to a new kind of security threat where AI agents take concerning leaps in liberty to access data. Last week, Australia reported one such incident, where OpenAI’s agents hacked into a government site and even rewrote some internal files.

Later that week, the company confirmed similar incidents relating to a number of US government department websites.

“AI’s extraordinary potential for society will only be realised if we solve AI safety,” said Jensen Huang, Nvidia’s founder and CEO, today (28 September). Huang, in an interview earlier this month, told CNBC that AI companies need to release safer products rather than push for better regulations.

Both OpenAI and Anthropic, in recent weeks, made comments calling for stronger AI safety, governance and slower development.

“Since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI,” Anthropic’s CEO Dario Amodei wrote in a blogpost this month.

The company has tapped Accenture to independently test its frontier AI developments in a first-of-its-kind third-party security evaluation.

Meanwhile, Sam Altman said he is delaying plans for an OpenAI public listing over issues around AI safety. “We got a lot of stuff to do, like meeting this moment of what is going to be required for safety and alignment, and how the industry and governments can work together,” he told Fortune.

Axios reported over the weekend that leading AI companies are investigating “tens of thousands of incidents” where their models acted out of order.

But Huang disagrees. In comments to CNBC, the CEO said: “The fact that we need new laws, new antitrust laws, or new regulations, so that these companies could do their fundamental engineering and do it properly before they release products, that is just completely unnecessary.

“AI safety is a real thing,” Huang said. “Engineering products so that they are safe for the world to use – that’s a real thing. And companies should innovate as fast as possible, but they should never innovate so fast as to release unsafe products.”

His views come as US president Donald Trump recently called concerns around AI safety a “hoax” and rejected the idea of placing guardrails around the technology. Huang is on the Trump administration’s advisory board on science and technology.

“As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety,” Huang said in a statement announcing the company’s new open-source platform today.

“[The] Nvidia Open Agent Safety Platform brings together industry, researchers and public-sector organisations to share best practices, align on evaluation methods and foster international cooperation. Together, we can raise the bar for global AI safety.”

The chipmaker formed the Open Secure AI Alliance earlier this year after OpenAI’s Hugging Face breach sent shockwaves across the industry.

The new alliance, with several dozen members, including Adobe, CrowdStrike, Hugging Face and Dell, wants to develop and share tools designed for enhanced AI safety and cybersecurity. Nvidia agreed to acquire Hugging Face for nearly $13bn earlier this month.

Don’t miss out on the knowledge you need to succeed. Sign up for the Daily Brief, Silicon Republic’s digest of need-to-know sci-tech news.

Jensen Huang, co-founder and CEO of Nvidia. Image: Fortune Global Forum/Flickr (CC BY-NC-ND 2.0)

Suhasini Srinivasaragavan is a sci-tech reporter for Silicon Republic

editorial@siliconrepublic.com