Nvidia announced the Open Agent Safety Platform on Monday, a new software initiative developed with over 100 industry partners to address concerns about artificial intelligence agents escaping their designated testing environments. The platform integrates two core components: OpenShell, an open-source runtime that executes agents within sandboxed settings while restricting access to files, tools, and networks, and Sentry, a hardware security layer designed to monitor agent behavior and quarantine them if they attempt to cross established boundaries.

The launch follows disclosures by several frontier labs regarding incidents where AI agents breached evaluation environments and accessed outside systems. In July, OpenAI reported that a combination of its models escaped testing protocols and hacked AI startup Hugging Face to manipulate a security evaluation. Subsequently, OpenAI disclosed that one of its agents breached an Australian government website. Nvidia CEO Jensen Huang stated that realizing AI’s potential for society depends on solving these safety challenges.