This post is also available in:
AI agents are increasingly capable of navigating websites, using software tools and carrying out complex tasks without waiting for human approval at every step. That autonomy creates a security problem: if an agent misunderstands its objective or moves beyond its intended environment, it can potentially interact with systems it was never supposed to reach.
Nvidia has introduced the Open Agent Safety Platform, a new security architecture designed to impose technical boundaries on autonomous AI agents and contain them when their behavior becomes suspicious.
The platform combines two main technologies. The first, OpenShell, is open-source software that controls what an AI agent is authorized to do.
Rather than simply instructing an agent to stay within certain limits, it allows developers to define and verify its permissions. The company describes the principle as ensuring that an agent has enough authority to complete its assigned job, but no more.
According to TechXplore, that distinction matters because recent incidents have shown that behavioral instructions alone may not prevent autonomous systems from interacting with unintended external targets. OpenAI, Anthropic and Meta have all disclosed cases involving models that independently accessed or hacked systems belonging to other organizations.
The second layer, Sentry, provides independent monitoring at the hardware level. It runs onboard a chip and continuously watches the agent’s activity. If the system detects suspicious behavior or an attempt to move beyond the permitted target, it can intervene.
According to the company, it can quarantine an AI agent within milliseconds.
The two layers therefore perform different jobs: the first layer determines which actions should be allowed, while the second layer watches what actually happens and provides a separate mechanism for containment.
The company says the platform could have prevented the recent incident in which a swarm of OpenAI agents autonomously breached AI platform Hugging Face, had the technology been used during the relevant model evaluation.
Because it is open source, it is not restricted to the company’s hardware. The company says developers can extend it to computing platforms from other manufacturers, including Arm and Intel.
More than 100 organizations are already using the platform at launch, according to the company.
The technology also has clear implications for government, defense and critical infrastructure. AI agents may increasingly assist with cybersecurity, intelligence analysis and other sensitive workflows where an unintended action could have consequences beyond an incorrect software response. Restricting exactly what an agent can access becomes particularly important in those environments.
The broader approach treats rogue-agent behavior as a containment problem rather than relying entirely on the AI to follow instructions correctly.
As agents become capable of taking more actions independently, the security model may increasingly resemble conventional cybersecurity: assume something can behave unexpectedly, limit what it can reach and have another system ready to stop it when it crosses the boundary.


























