NVIDIA launched its Open Agent Safety Platform on 28 September: open software plus a reference hardware design meant to keep AI agents inside the limits their operators set. It follows reports from several frontier labs of agents breaking out of the test environments meant to contain them.
There are two parts. OpenShell, open-source runtime software that is now broadly available, runs each agent in a sandbox, traces its actions and enforces a policy on which files, networks, tools and credentials it may use. Sentry, a reference design that runs on NVIDIA’s BlueField-4 data processing units, monitors agent activity from hardware the agent cannot see and, NVIDIA says, can quarantine an agent that steps outside its boundary within milliseconds.
NVIDIA names Anthropic, Microsoft, Hugging Face, Palantir, SAP and SpaceXAI among the platform’s backers. Anthropic separately said it is pairing OpenShell with Claude Managed Agents, its service in which credentials sit in a vault the agent never sees and each agent’s actions are logged.
Why it matters: the premise is that an agent cannot be trusted to police itself, so the controls move outside it. Sentry matters most to data centres already on NVIDIA hardware; NVIDIA gives no pricing and says some features will be offered on a when-and-if-available basis.
