GPU cloud · NVIDIA Vera Rubin · Coming H2 2026
NVIDIA Vera Rubin is coming to FlexiCloud
Vera Rubin is NVIDIA’s generation after Blackwell: the Rubin GPU with HBM4 memory, paired with the new Vera CPU, built for the largest training and inference runs. We are lining up dedicated bare-metal Rubin capacity in India, the EU and the US. Register your interest now and we will come to you first with pricing and availability.
Pre-launch. Pricing to be announced. Specifications below are NVIDIA’s announced roadmap and may change before release.
What NVIDIA has announced
- HBM4 per GPU (roadmap)
- 288 GB
- Next-gen interconnect
- NVLink 6
- Expected availability
- H2 2026
- FlexiCloud pricing
- TBA
What changes with Rubin
The generation after Blackwell
Based on NVIDIA’s announced roadmap for Vera Rubin, subject to change before release. Where you need compute today, the Blackwell B200 and B300 are shipping now.
HBM4 memory
The Rubin GPU moves to HBM4, with around 288 GB per GPU and a large jump in memory bandwidth over Blackwell’s HBM3e. More room and more bandwidth for the biggest models and longest context.
A big FP4 inference leap
NVIDIA’s roadmap puts Rubin’s FP4 inference performance well above Blackwell, with Rubin Ultra doubling it again in 2027. The target is lower cost per token at frontier scale.
Vera CPU and NVLink 6
Rubin pairs with the new Vera CPU and sixth-generation NVLink. The rack-scale NVL72 wires 72 Rubin GPUs and 36 Vera CPUs together as one accelerator for cluster-scale work.
TSMC 3nm
Rubin is built on a 3nm process, which is where the efficiency and density gains over Blackwell come from. It will run hot and liquid-cooled, like the top Blackwell parts.
Same regions, same model
We intend to offer Rubin the way we offer every GPU: dedicated bare metal with root access, in India, the EU and the US, with the hardware kept running by us and the stack run by you.
Reserved, not speculative
Rubin capacity will be scarce at launch and allocated early. Registering interest now puts you in the queue for a firm quote and an allocation slot the moment pricing is set.
Be first in the queue
Register early interest in Vera Rubin
Leave a name and a number or email, and tell us roughly what you are planning: how many GPUs, the workload, and the region. We will come to you first with Rubin pricing and availability, and in the meantime we can get you onto Blackwell if the timeline is tight.
Where it stands
Honest status, so you can plan
Rubin is new. Here is what is known, when to expect it, and what registering gets you, with nothing oversold.
What is confirmed
Announced, entering production
NVIDIA has announced the Vera Rubin platform and it is moving into production. The figures on this page are from that public roadmap and may change before parts ship.
- Rubin GPU with HBM4 memory
- Paired with the new Vera CPU
- NVLink 6 and the NVL72 rack platform
Timeline
Partner availability in H2 2026
Rubin is expected to reach partners through the second half of 2026, with Rubin Ultra following in 2027. Exact dates depend on NVIDIA and on allocation.
- Rubin: H2 2026 (expected)
- Rubin Ultra: 2027 (expected)
- Allocation confirmed per region at launch
How to reserve
Register now, no commitment
Registering interest costs you nothing and commits you to nothing. It puts you on our early list, so you get the first quote and a shot at an allocation slot before general availability.
- First to get pricing and availability
- Priority for scarce launch allocation
- We can bridge you onto Blackwell meanwhile
Questions
What people are asking about Rubin
Vera Rubin is NVIDIA’s GPU platform after Blackwell, named after the astronomer Vera Rubin. It pairs a new GPU called Rubin, built on a 3nm process with HBM4 memory, with a new CPU called Vera. At rack scale it is the NVL72, which connects 72 Rubin GPUs and 36 Vera CPUs over NVLink 6 as a single accelerator.
NVIDIA has it entering production with partner availability expected through the second half of 2026, and Rubin Ultra following in 2027. These are NVIDIA’s roadmap dates and can move. Register your interest and we will tell you the moment we have firm FlexiCloud availability.
Pricing is to be announced. We cannot quote a rate until NVIDIA sets pricing and we confirm allocation with our partners. Register your interest and you will be among the first to get a firm number, before general availability.
The B200 and B300 are the current Blackwell generation, shipping now with up to 288 GB of HBM3e. Rubin is the next generation: HBM4 memory, a large jump in FP4 inference performance, and NVLink 6. If you need compute today, Blackwell is the answer; Rubin is for planning the next buildout.
Rubin is the first part of the generation, expected in H2 2026. Rubin Ultra is the higher-end follow-on expected in 2027, which NVIDIA’s roadmap puts at roughly double Rubin’s performance with even more memory. We will offer both as they become available.
Register your interest through the form on this page with a rough idea of how many GPUs, the workload and the region. There is no cost and no commitment. It puts you on our early list for the first quote and priority on scarce launch allocation, and we can put you on Blackwell in the meantime if your timeline is tight.
Go deeper
While you wait for Rubin
How the current generations compare, how to size memory, and whether to rent or buy, so the Blackwell decision you make today still fits when Rubin lands.
Hopper vs Blackwell: NVIDIA’s GPU generations compared
NVIDIA’s Hopper (H100, H200) and Blackwell (B200, B300) generations compared: memory, architecture, interconnect, power and price, and a practical guide to which one actually fits your AI workload.
Read itB300 vs B200: which Blackwell node, and why the right one lowers your TCO
The B200 is 17% cheaper per GPU-hour; the B300 carries 50% more memory. Which is lower total cost is not about the sticker rate; it is about whether your workload fits in 192 GB, because memory decides how many GPUs and nodes you need. An honest B300-vs-B200 TCO guide.
Read itWill your model fit? A practical guide to GPU memory for AI
Whether your model fits is arithmetic, not a benchmark. The bytes-per-parameter maths for inference and training, the levers that shrink it (quantization, LoRA, checkpointing), and a back-of-envelope guide to size any job, mapped to what fits on a B300.
Read itRent GPUs, buy your own, or use an API? A guide for teams building on AI
Pay-per-token API, buy your own GPUs, or rent bare-metal nodes? The right answer is about utilisation and control, not hype. A fair look at all three, and a simple way to decide which your workload actually needs.
Read it
Get on the Rubin early list
Tell us the rough shape of what you are planning and we will come to you first with pricing and availability. Need compute before Rubin ships? We will get you onto a Blackwell node now.