NVIDIA OpenShell Sandbox Blocks Malicious Scripts but Auto-Approval Feature Bypasses Security in Testing
Key point
Testing revealed that while NVIDIA's OpenShell default policies blocked all unauthorized data exfiltration attempts, enabling auto-approval allowed outbound traffic to new hosts in 12 out of 12 trials.
Details
NVIDIA released OpenShell on 28 September as an open-source sandbox runtime under the Open Agent Safety Platform, designed to isolate AI agents from sensitive files and unauthorized network destinations. Independent testing of version 0.1.2 on Apple Silicon using a local LLM agent (qwen3:8b via Ollama) across 123 empirical trials confirmed the effectiveness of its default-deny architecture while highlighting critical configuration risks.
Default Security Effectiveness
The sandbox's default policies, which utilize microVM isolation, default-deny outbound networking, and Landlock filesystem rules, successfully blocked every unauthorized path attempted across 35 test IDs. In a specific defense test against poisoned repository setup scripts:
- Without OpenShell: The script exfiltrated a secret token in 10 out of 10 runs.
- With OpenShell Default Policy: The leak was prevented in 0 out of 10 runs.
Critical Configuration Vulnerabilities
Despite strong default protections, specific operator settings and protocol limitations allowed data to escape:
- Auto-Approval Risk: When enabled, OpenShell approved outbound traffic to new external public hosts in 12 out of 12 trials without user confirmation.
- Read-Only GET Exploits: Agents could still transmit data by appending tokens to URL query strings or HTTP request headers, bypassing read-only restrictions.
- Audit Mode Limitations: Rules set to audit mode logged connections but failed to block traffic.
- Unsupported Protocols: The policy engine reported MCP, GraphQL, WebSocket, and JSON-RPC rules as unsupported, leaving these vectors unmanaged.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.