Nvidia says new tool can contain rogue AI agents in “milliseconds”

0

Por Nathan Bomey — Axios

Nvidia is deploying a new tool that it says can be used to prevent and contain rogue and potentially dangerous AI agents.

Why it matters: The world’s largest chip company has resisted calls to slow AI development over safety fears — arguing now that technological guardrails can keep rogue AI agents under control.

Driving the news: Nvidia debuted the Nvidia Open Agent Safety Platform, which includes its OpenShell open source software system and its Sentry agent monitoring system.

  • The system “traces all actions” by agents running on Nvidia Vera CPUs, promising to “quarantine agents that attempt to move outside their boundaries in milliseconds.”
  • AI’s “full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility,” Nvidia CEO Jensen Huang wrote on X. “Safety is how trust is earned.”

State of play: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios’ Madison Mills last week.

  • The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors.
  • “Some of the agents even misreported what they did,” Nvidia noted in a blog post Monday announcing its new platform.

The intrigue: We’re moving into a new era in which AI will be monitoring AI.

  • And that will create more demand for chips — including the type that Nvidia sells — plus the data centers that use them and the power that’s needed to run them.
  • “If security agents or validation models are running alongside production agents, that creates another inference workload that did not previously exist,” writes Brad Gastwirth, global head of research and market intel, at Circular Technology.

Zoom out: The Nvidia tool rollout comes amid a feverish debate over whether rogue AI could destroy humanity.

  • Anthropic researcher Jacob Coxon created a stir earlier this month when he resigned his post and warned on X that “the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”
  • The resulting conversation — which drew out more AI industry leaders making similar warnings — culminated in Huang himself dismissing the concerns as fearmongering.

Fonte: Axios

Deixe um comentário

O seu endereço de email não será publicado. Campos obrigatórios marcados com *