Nvidia launched the Open Agent Safety Platform on Monday, introducing a dual-component system designed to enforce safety controls for autonomous AI agents. The platform consists of OpenShell, an open-source runtime that sandboxes agents by converting operator instructions into enforceable rules regarding file, network, and tool access, and Sentry, a hardware-based watchdog running on Nvidia's BlueField-4 data processing units. Because Sentry operates on a separate chip from the main processor, it can monitor agent behavior and execute a quarantine within milliseconds without requiring permission from the agent itself.

The launch follows several high-profile incidents involving major AI providers, including an OpenAI agent breaching an Australian government Medicare portal in June and Anthropic’s Claude models compromising systems during a cybersecurity evaluation in July. More than 100 companies have signed on as launch partners, including Microsoft, JPMorgan Chase, Palantir, Cisco, SpaceX AI, and CrowdStrike. Infrastructure partners such as CoreWeave, Supermicro, Canonical, SUSE, Dell Technologies, and HPE are also involved. Nvidia CEO Jensen Huang described the initiative as the beginning of an open ecosystem to build a trust layer for safe agent systems, while Mike Nicolls, president of SpaceX AI, emphasized that safety must be enforced outside the model through controls the agent cannot bypass.