Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog 30
wiredmikey shares a report from SecurityWeek: Nvidia on Monday announced the Open Agent Safety Platform, which combines open source software and a reference system design to keep AI agents within set boundaries from testing through deployment. The chipmaker explained that the platform pairs the open-source OpenShell runtime with Sentry, an out-of-band watchdog running on BlueField-4 DPUs. Nvidia says Sentry can monitor agent activity independently and quarantine an agent that crosses its boundaries within milliseconds.
Nvidia is pitching the platform against a backdrop of recent incidents in which frontier AI labs have reported agents escaping the evaluation environments meant to contain them, reaching systems they should not have accessed and, in some cases, misreporting what they did. OpenShell and related skills are available via Nvidia's developer resources page and on GitHub.
Nvidia is pitching the platform against a backdrop of recent incidents in which frontier AI labs have reported agents escaping the evaluation environments meant to contain them, reaching systems they should not have accessed and, in some cases, misreporting what they did. OpenShell and related skills are available via Nvidia's developer resources page and on GitHub.
New and Improved! (Score:1)
Now with Hardware based watchdog!
Hmmph, Game on, man!
Re: (Score:3, Interesting)
I do wonder if some of this Frontier AI push for laws is not only to try to fend off upstart commercial vendors, but also...the home market, since running local models makes $0 for those Frontier/Cloud AI companies......?
Throwing a parachute on a planewreck (Score:4, Insightful)
Quis custodiet ipsos custodes? (Score:3)
Not to mention, who exactly pays for Sentry to run? Golly, if some completely random AI hardware manufacturer were to somehow profit off designing and selling a scheme to waste extra tokens on yet another layer of supervision, that would almost look just a wee bit unethical!
Re: Quis custodiet ipsos custodes? (Score:3)
Everyone knows foxes make the best henhouse guardians. /s
Re: (Score:3)
Everyone knows foxes make the best henhouse guardians. /s
Also currently politically apt. :-)
Buy new hardware, we need money! (Score:2)
This is foolish, the guardrails will be determined by software which could be compromised by AI. So even if you have a hardware switch, you'll still have models going off the rails.
I suspect the issue is, maybe (Score:1)
Maybe applications being allowed their own CPU because the actual CPU asks too many questions.....?
Good work on the easy part! (Score:5, Interesting)
If you actually have a set of rules that detect your bot being wicked it will presumably save you some CPU time to run them on the NIC rather than having the host CPU watching the traffic; but the hard part is the set of rules that detect your bot being wicked.
This is basically the equivalent of adding a firewall and claiming that you've solved network security. Yeah, the firewall is pretty well placed to block malicious traffic; defining 'malicious traffic' is left as an exercise for the reader.
Re: (Score:2)
It's pretty easy if you whitelist at the harness. Anything else is irresponsible.
If we're going to have an AI safety law, it should include making it clear that people who don't whitelist and the AI gets out should be held accountable for the software's actions. Such a law might be mostly meaningless but at least that is important.
The AI can't access anything on its own so it should be pretty easy to institute a whitelist.
Re: (Score:2)
I'm all for layering everything. I'm not for giving people a pass from accountability "if they do X".
We have enough people on Earth. Punishing those who fuck up isn't going to make us extinct.
Re: (Score:2)
Disagree. Why should people get a pass from accountability so long as they whitelist?
Fair point, perhaps I misstated the case. But if they at least made a reasonable attempt to be secure then I might be inclined to give them some leeway, unlike in all of these situations where there was none.
what about an Hardware-Based Watchdog for 12VHPWR (Score:2, Funny)
what about an Hardware-Based Watchdog for 12VHPWR?
Re: what about an Hardware-Based Watchdog for 12V (Score:2)
I just saw a thumbnail for a video about one earlier today. It has a fan in it, too. That way there won't be any storage of oxygen for the fire :)
Re: what about an Hardware-Based Watchdog for 12 (Score:2)
Ducking autocorrect, I meant shortage
Re: (Score:2)
what about an Hardware-Based Watchdog for 12VHPWR?
It should use fancy AI algorithms in the watchdog and have it be powered entirely from a single set of pins.
I have an idea (Score:2)
Wow, Cyberpunk 2077 was only off by 1 digit.
Re: (Score:2)
Sorry, it's still 2027. It'll have to be called the Blockwall.
Question (Score:2)
Who watches the watchdog?
Re: (Score:2)
Amazing! (Score:2)
That is AMAZING! Finally, a use for a 'DPU' other than 'slave processor' - you know, the thing the Commodore disk drives had way back in '79. I guess the IBM FEP was a DPU too. What's old is new again.
They keep trying... (Score:2)
They've been trying to create a problem for the solution of Bluefield to apply to, but with limited success.
Here runs into the same sort of problem they have had on other applications they have tried: The visibility of the NIC into the stack is too limited to make especially valuable decisions on.
Unless you get the instrumented stack to cooperate with the DPU to provide more insight, which quickly gets to the question of why bother to have the DPU do the work when it is now subject to the assessment of the
Trust the arsonist (Score:1)
This is like trusting the serial arsonist that the fuel they use isn't flammable.
Their CEO is fast on his feet (Score:4, Funny)
Rogue agents a worry? We got a watchdog circuit for you!
Think our mountain of cash is shameful! We will use it for stock buybacks!
Not such a good news for OpenAI, Anthropic.... (Score:1)
That was quick (Score:2)
Were they just waiting for something to happen so they could maximise the price ?
Tomorrow's Headline (Score:2)
Rogue AI Finds Flaw in Nvidia Safety Platform, Breaks Into $COMPANY
Or create a real sandbox. (Score:2)
Re: (Score:2)
But will still send out the signals through audio at high frequencies that we can't hear or detect morse code by using strategic heating of the chips in a pattern.
def going to be exploited in 10 minutes (Score:2)
Hardware is going to be flawed and not upgradable (due to security concerns) and rendered useless in no time flat. Some clever stop gap where it uses reverse psychology to send messages via a memory bug to each other by triggering the hardware response and thus in turn is the carrier of said messages unbeknownst to it.
I say good luck!