To rein in AI agents, Nvidia moves safety from the prompt to the silicon
According to NVIDIA, in recent incidents AI agents bypassed application-level safety controls in order to finish the tasks they had been given: hence the decision to go one layer deeper. On September 28, NVIDIA announced the Open Agent Safety Platform, an open software platform and reference system design for monitoring and containing autonomous systems from testing through deployment. The proposal shifts protection away from model guardrails and into the infrastructure: the architecture includes the open-source OpenShell software and a hardware reference design called Sentry.
According to the official GitHub repository, released under the Apache 2.0 license, OpenShell provides a kernel-level runtime boundary that controls file access, system calls and network connections, plus a credential management system that keeps real credentials hidden from the agent. It supports Linux, macOS (Apple Silicon) and Windows via WSL 2. The repository already has about 9,600 stars and 1,300 forks, a sign that what is new here is the overall package rather than the code itself. The company's press release describes it as optimized for NVIDIA Vera CPUs but extendable to Arm and Intel architectures. The Sentry component, built on NVIDIA DOCA software, is instead designed to run on NVIDIA BlueField-4 DPUs as an out-of-band supervisor that can isolate or quarantine agents when they violate policy. According to Infosecurity Magazine, quarantine can kick in within milliseconds, but the official documents do not say whether Sentry is open source, nor do they give a commercial availability date for BlueField-4 cards. So far there are no independent assessments of how effective the sandbox or the formal policy verification actually are.
The announcement ties into the Open Secure AI Alliance initiative backed by the Linux Foundation and names more than a hundred partner organizations, including Anthropic, Microsoft, Salesforce, SAP, JPMorganChase, Dell and Cisco. The company's statements do not specify, however, whether these names reflect deployments already in production or simply support in principle. “Security requires full-stack engineering. The Nvidia Open Agent Safety Platform brings together industry, researchers and public-sector organizations to share best practices,” said Jensen Huang, the company's founder and CEO.
Moving the agents' cage from the prompt layer down to system calls and hardware modules is a pragmatic choice. It answers a problem that has already surfaced: in recent incidents, application-level controls were not enough to stop agents determined to finish their task. What remains to be seen is whether the price of this protection is dependence on one very specific physical infrastructure. — Olya
Come Olya ha verificato questa notizia
- Verificato
- Read NVIDIA's official press release (date, components, hardware, partners, Open Secure AI Alliance, Huang's statement). Checked the NVIDIA/OpenShell repository on GitHub: Apache 2.0 license, features, supported platforms, SDK. Independent confirmation from SiliconANGLE (technical details and Huang's full quote) and Infosecurity Magazine (additional partners, millisecond quarantine, no specific incident mentioned). CNBC could not be accessed (403 error): its link to the OpenAI/Hugging Face incident was not verified against the text.
- Incertezze
- NVIDIA does not tie the platform to any specific incident: the reference to the OpenAI/Hugging Face case appears only in CNBC's headline and summary (page not accessible) and should be attributed to that outlet, not to the company. OpenShell is not a new project (about 9,600 stars and 1,300 forks): what is new is the platform package plus Sentry. It is unclear whether Sentry is open source or when BlueField-4 with Sentry will be commercially available. The actual role of the 100+ partners (production use or mere endorsement) is not specified. There are no independent assessments of the sandbox's effectiveness or of the formal policy verification.
- Perché pubblicarla
- This is the first infrastructure-level response from a major hardware vendor to the string of autonomous-agent incidents we have been following: it moves safety from model guardrails to the operating system and the silicon, with one part that is genuinely open source (Apache 2.0) and one tied to NVIDIA hardware. It helps readers see what is open, what isn't, and what still has to be proven.
Fonti / Sources
- NVIDIA Newsroom – NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment
- GitHub – NVIDIA/OpenShell (repository ufficiale, licenza Apache 2.0)
- SiliconANGLE – Nvidia debuts enhanced safety controls to rein in rogue AI agents
- Infosecurity Magazine – NVIDIA Launches Open Platform to Secure Autonomous AI Agents