CoreWeave brings NVIDIA Vera Rubin NVL72 and Vera CPU to production with Cognition as first customer
Key point
Cognition benchmarks show Vera Rubin NVL72 delivers up to 4.8x higher token throughput for SWE-2 inference compared to GB200 NVL72.
Details
CoreWeave announced the availability of NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet networking on its cloud, marking one of the first deliveries of this platform to customers. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on the new infrastructure.
Performance Gains for Agentic Coding
Cognition benchmarked Vera Rubin’s inference performance against a GB200 NVL72 baseline using real-world software engineering tasks from FrontierCode. The tests revealed that Vera Rubin NVL72 delivered up to a 4.8x increase in total token throughput for SWE-2 inference workloads. These gains enable faster real-time code generation and more responsive multistep reasoning for Devin.
NVIDIA Vera CPU for Agent Sandboxes
CoreWeave also announced that NVIDIA Vera, the first CPU built specifically for AI agents, will be available on its cloud. A single Vera rack contains 128 CPUs and 11,264 cores, supporting more than 11,000 concurrent isolated environments. In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on Vera CPUs and a 1.7x performance gain on Terminal-Bench across all passing tasks.
CoreWeave Forge Unifies the AI Loop
To close the gap between production and training, CoreWeave launched CoreWeave Forge, a connected environment integrating Weights & Biases, OpenPipe post-training expertise, and the marimo notebook project. Key components include:
- CoreWeave ARIA: Now generally available, it analyzes runs, proposes experiments, and recommends code changes.
- CoreWeave Agent Lens: A new service that improves failure detection by 20% and fixes issues at half the cost.
- CoreWeave Sandboxes: Now generally available, these provide isolated CPU or GPU execution environments for agents and RL runs.
- Serverless RL: Trains 1.4x faster at 40% lower cost than self-managed setups.
Early adopters of Forge include Canva, Capital One, and MasterClass. The platform supports NVIDIA Nemotron open models for customizing reasoning and multimodal workflows.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.