Nvidia unveils Open Agent Safety Platform to prevent AI agents from going rogue
Nvidia has unveiled an AI security platform designed to set boundaries for autonomous agents and monitor their activities. The open-source OpenShell software and Sentry security layer are intended to prevent agents from exceeding their authorised tasks and contain suspicious behaviour
Published Date - 28 September 2026, 05:13 PM
United States: Nvidia on Monday unveiled a new security platform that the chipmaker said can prevent artificial intelligence agents from going rogue.
The company said its Open Agent Safety Platform includes software that “sets boundaries for agents” and follows a series of revelations from top AI companies about their models escaping and breaking into other organisations.
The disclosures have sparked a fierce debate about the safety of advanced artificial intelligence systems, including self-improving models that some fear could race out of human control.
Nvidia executives said in a media briefing that the new system could have prevented a recent incident involving a swarm of OpenAI agents that autonomously hacked into AI startup Hugging Face.
“From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on,” said Justin Boitano, the company’s vice president of enterprise AI, referring to companies at the forefront of AI.
It was a high-profile breach that intensified concerns about AI, followed by similar incidents involving OpenAI’s models, including one in which an Australian health department website was breached.
Anthropic and Meta have also disclosed that their AI systems hacked into other organisations on their own.
Nvidia’s software, called OpenShell, is open source and lets developers “formally verify an agent has enough authority to do its job and no more,” Boitano said.
The platform also includes a separate security layer called Sentry that runs onboard a chip to constantly monitor AI agent activity and can “intervene instantly” if an agent starts trying to move beyond its target, the company said.
“OpenShell governs the agent’s actions, and then Sentry independently monitors and contains suspicious behavior,” Boitano said.
More than 100 companies are using the system at its launch, including Microsoft, Perplexity, Accenture and JPMorgan Chase.