Cognition became the first customer running production workloads on the new architecture.
CoreWeave announced the production availability of NVIDIA Vera Rubin NVL72 systems on September 30, 2026, during its Fully Connected conference in San Francisco. Applied AI lab Cognition is the first customer to run production workloads on the new architecture. Cognition uses the cluster to handle training, reinforcement learning, and inference for its Devin coding assistant.
Cognition stood up a Vera Rubin NVL72 cluster with CoreWeave in early September 2026. In benchmarks against a GB200 NVL72 baseline, Cognition reported up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. The team also recorded a 3.8 times gain in output token throughput for reinforcement learning workloads. Cognition scaled from bridge capacity to thousands of GPUs on CoreWeave in under nine months.
CoreWeave also published performance data showing 10 times the token throughput per megawatt over NVIDIA GB200 NVL72 systems on the DeepSeek R1 reasoning model at matched interactivity. The company collaborated with Dell Technologies on deploying Dell PowerRack systems for GB200 and GB300 NVL72, and is among the first to bring Vera Rubin to market.
The deployment expands a partnership between CoreWeave and NVIDIA that began in 2017 with the Volta generation, which remains in commercial service on CoreWeave Cloud. CoreWeave completed its public listing on Nasdaq in March 2025 and serves nine of the top ten foundation model providers.
Newsletter
Markets in your inbox, weekly
Latin America-focused analysis, investment themes and the week in finance.
Keep reading