
CoreWeave is opening Nvidia’s new Vera Rubin systems to its first customers. Cognition already uses them for its coding agent Devin and reports substantially higher throughput. The availability announced on September 30 shows the next hardware generation moving into everyday AI services. For users, the key question is how much of that acceleration reaches a completed task.
Key takeaways
- Vera Rubin NVL72 has limited availability on CoreWeave; Cognition already runs production workloads on the systems.
- For SWE-2, Cognition reports up to 4.8 times the total token throughput of GB200 NVL72. This measures computing performance, rather than the quality of completed software.
- An NVL72 system connects 72 Rubin GPUs with 36 Vera CPUs. The two handle different types of work.
- Interested teams can inquire about capacity or request an ARENA Pass for their own workloads. This does not establish a general pricing or availability commitment.
What is actually available now
The news marks a specific step: according to CoreWeave, customer workloads run on Vera Rubin NVL72, rather than the platform appearing solely in future expansion plans. The provider explicitly describes limited availability. Anyone planning a computing project still needs to establish when capacity will be available and how much they can obtain. A product announcement does not mean every customer can freely book an allocation.
Cognition is the first named production customer. The company behind Devin uses CoreWeave for training, reinforcement learning, and production inference. Reinforcement learning evaluates attempts at a solution so subsequent attempts can deliver better results. In production, an already trained model computes responses, a process known as inference. A coding agent also needs tools, executable environments, and feedback from tests.
This distinction helps explain the significance of the hardware. A chat can end after a single response. An agent asked to fix an application bug might first have to read files, write a change, run a test, and respond to the result. That is an example of a possible workflow, rather than a promise about how Devin handles every task. Acceleration affects individual parts of that process.
What the reported multiplier measures
Cognition’s engineers compared SWE-2 inference on Vera Rubin NVL72 with a GB200 NVL72 baseline. The published chart describes up to 4.8 times higher total token throughput per GPU at matched interactivity. Tokens are the units of text that a model processes or generates. The result is a customer-specific benchmark published by CoreWeave; it does not replace an independent comparison across all models and applications.
Higher throughput can allow a service to handle more work concurrently without slowing generation in each session. That is valuable to the operator. Someone sitting at a screen waiting for a bug fix is asking a different question: how long does the entire task take? That includes waiting for tools, repeated attempts, and checking the result. Higher token throughput does not directly determine that total duration.
It also does not establish that the agent solves more tasks correctly. Faster text computation can deliver the same answer sooner or enable more attempts; their usefulness needs to be measured separately. Nor has a corresponding reduction in subscription prices been announced. Hardware costs, utilization, and a service’s pricing decisions are separate factors. Reading the multiplier as a universal acceleration of one’s own programming work skips over those distinctions.
Why GPUs alone are not enough
Nvidia’s product documentation lists 72 Rubin GPUs and 36 Vera CPUs for NVL72. GPUs handle the highly parallel calculations inside the model. CPUs also perform work outside it, including code execution, data processing, and coordinating tools. Faster responses offer limited help if the environment where an agent tries out its suggestions repeatedly becomes the bottleneck.
CoreWeave also announced Vera as a standalone CPU offering. This planned service needs to be distinguished from the reported production use of complete NVL72 systems. The general Vera product page and CoreWeave’s specific expansion announcement describe different offerings. An evaluation therefore depends on more than a processor name: it needs to establish which system a team can actually order and use for its work.
At the same event, CoreWeave introduced Forge. The development environment brings together Weights & Biases, OpenPipe, and marimo, among other components, with the company’s own services. Experiments, evaluations, and model versions are intended to sit within a connected workflow. The underlying idea makes sense for developers: a production failure should become a reproducible test case against which the next change can be checked.
A fast execution environment does not automatically determine which data an agent may access. Our article on Nvidia’s OpenShell and Sentry explains the separate task of limiting tools and permissions. Computing performance and contained execution complement each other; the presence of new hardware alone says nothing about the permissions assigned to a particular application.
Who should run a comparison of their own
The immediate audience consists of teams operating models themselves or running many agent tasks concurrently. CoreWeave points to capacity planning briefings and the ARENA Pass, which lets interested teams try their own workloads on Vera Rubin NVL72. A useful comparison would involve a representative workflow with the same model, a similar number of concurrent tasks, and a clear success criterion. It could then establish the cost and duration of a correctly completed task.
For users of a finished AI service, the advance is indirect: its operator gains new options for expanding infrastructure. Whether those lead to shorter waits, additional capacity, or different prices depends on the individual product. The next relevant measurement is therefore more than a count of generated tokens. It is the path from a request to a usable result, including the steps that take place outside the language model.
Sources
- CoreWeave: Cognition und SWE-2-Benchmarks auf Vera Rubin NVL72
- CoreWeave: Produktionsverfügbarkeit, 30. September 2026
- NVIDIA: Vera Rubin NVL72, technische Produktunterlagen
- NVIDIA: Vera CPU, Aufgaben und Architektur
- NVIDIA: CoreWeave-Ausbau und angekündigtes Vera-Angebot
- CoreWeave: Forge-Ankündigung
- Data Center Dynamics: Bericht zur CoreWeave-Verfügbarkeit

