CoreWeave Opens Limited Access to Nvidia’s Vera Rubin, With Cognition First in Production
Cognition’s customer benchmarks show higher token throughput at matched responsiveness, but they do not measure how much faster Devin completes coding assignments.
CoreWeave’s September 30, 2026, limited rollout puts Nvidia Vera Rubin NVL72 into production for select customers, with Cognition first rather than merely announcing a future deployment. Cognition reports up to 4.8× inference token throughput per GPU and 3.8× reinforcement-learning output throughput versus GB200 NVL72, with responsiveness held constant; these are customer-run benchmarks, not independent validation or evidence of faster completed coding tasks. Hundreds of Rubin GPUs are deployed across multiple regions, but capacity is not generally open, so prospective users must assess workload fit through CoreWeave onboarding or testing.
01
Each Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs, linked by NVLink 6 and operated as one logical GPU.
02
The 3.8× reinforcement-learning result measures generated output throughput, a different metric from the 4.8× inference total-token result.
03
Cognition uses CoreWeave Inference for global agent queries and works with CoreWeave engineers on caching and serving configurations.
Select CoreWeave customers can now run production AI workloads on Nvidia’s Vera Rubin NVL72 systems. CoreWeave opened limited availability on September 30, 2026, naming Cognition as its first production customer. The maker of the Devin coding agent reports up to 4.8 times the inference token throughput per GPU compared with GB200 NVL72—a capacity gain measured at matched responsiveness.
CoreWeave says hundreds of Rubin GPUs are deployed across multiple regions, and it has begun onboarding production workloads. This is access for select customers, not a general opening of unlimited capacity. The announcement pairs that availability with a named customer already using the hardware, rather than presenting the platform solely as a future deployment.
Cognition’s engineering team compared Vera Rubin NVL72 with previous-generation GB200 NVL72 systems on CoreWeave. The inference test used SWE-2, Cognition’s latest generation of autonomous software engineer. CoreWeave describes the benchmarks as independently conducted by Cognition, but the results are customer claims published by the cloud provider, not an outside validation of either platform.
The comparison holds interactivity constant: the charts use output tokens per second per user as the responsiveness measure. That qualification matters because the claim concerns how much processing each GPU can sustain at the same interactive pace. It is not simply a claim that one user receives every response 4.8 times faster.
Cognition’s reported per-GPU gains
Up to 4.8×SWE-2 inference
Total token throughput per GPU versus GB200 NVL72 at matched interactivity, according to Cognition.
3.8×Reinforcement learning
Output token throughput per GPU versus GB200 NVL72 at matched interactivity, according to Cognition.
The two gains measure different things. Inference counts total token throughput; the reinforcement-learning result counts generated output. Neither figure directly measures completed coding assignments or code quality. They establish Cognition’s reported processing advantage for these workloads, not a single multiplier that applies to every model, training task or software project.
Cognition’s case for the upgrade rests on the repeated steps inside an agent’s work. Devin can work across entire code repositories, moving through reasoning, code generation, debugging, execution, evaluation and revision. Cognition research executive Silas Alberti says each step waits on the previous one, so faster processing can accumulate across a long assignment.
Production serving: Cognition uses CoreWeave Inference for real-time agent queries globally.
Performance tuning: Cognition works with CoreWeave engineers on memory-cache management, serving parameters and runtime configurations.
Experiment tracking: Cognition uses Weights & Biases Models to monitor and evaluate training experiments.
The training result addresses a separate part of that cycle. Cognition uses reinforcement learning—post-training built around trial-and-error reasoning—to improve Devin. Generating those trial sequences requires large volumes of output. CoreWeave argues that higher output throughput lets Cognition’s researchers iterate on model weights and move updates into production faster; it does not supply a measured reduction in that update timetable.
Vera Rubin NVL72 is a liquid-cooled rack-scale platform combining 72 Rubin GPUs and 36 Vera CPUs. Nvidia’s NVLink 6 connects the GPUs, and CoreWeave describes the system as engineered to operate as one logical GPU. The deployment therefore brings a coordinated computing system into the cloud, rather than offering the new processor in isolation.
CoreWeave says Cognition began running production workloads within days of rack handover, without doing custom infrastructure setup itself. For other prospective customers, the next step is a capacity-planning briefing covering onboarding timelines and workload fit. CoreWeave also offers an ARENA Pass to test a customer’s own workload on Vera Rubin NVL72—an opportunity to assess whether Cognition’s reported advantage carries over to a different job.
Sources
coreweave.comFirst Vera Rubin NVL72 Customer Sees 4.8x Throughput | CoreWeave
Reader comments
Newest comments first. Replies stay oldest first.