CoreWeave announced at its Fully Connected event in San Francisco on 30 September that Nvidia’s Vera Rubin NVL72 rack systems are available to customers, with Spectrum-X Ethernet networking. The first named production customer is Cognition, maker of the Devin software-engineering agent.
Cognition says it benchmarked Vera Rubin NVL72 against the previous GB200 NVL72 on real software-engineering tasks and saw “up to a 4.8x increase in total token throughput” for inference. That is the customer’s own figure on its own workload, not an independent benchmark. Its Silas Alberti put the reason plainly: in agentic coding, “cost per token decides what we can ship”.
Nvidia’s Vera CPU, which it describes as its first processor built for AI agents, is also coming to CoreWeave, with 128 CPUs and 11,264 cores in a rack, enough for over 11,000 concurrent agent sandboxes. CoreWeave also launched Forge, combining its Weights & Biases, OpenPipe and marimo products for training and evaluating agents.
Why it matters: agents run long, run in parallel and need somewhere safe to execute code, and the hardware is now being designed around that.
Source: Nvidia’s announcement.


