Skip to content

Renting a multi-node B300 cluster: what to check before you commit

28 September 2026The FlexiCloud Team
Renting a multi-node B300 cluster: what to check before you commit

A single 8-GPU B300 node is already a lot of compute. Before you rent a cluster of them, it is worth being sure you actually need one, and then knowing what you are really buying, because a cluster is far more than a pile of GPUs. This is the multi-node companion to our B300 node page: same honesty, harder problem. Once you have decided to rent, pair it with the questions to ask any GPU provider before you sign.

First, do you actually need multi-node?

One 8-GPU node pools a large amount of HBM across its GPUs, so a surprising number of models train and serve on a single node. You only need several nodes when the model plus its KV cache will not fit in one node, or when the training time on a single node is simply too long. If it fits, one node is cheaper and simpler, and the cheapest setup that avoids sharding across nodes usually wins. We worked through the memory maths in will your model fit. From here, assume you have decided you genuinely need a cluster.

A cluster is far more than the GPUs

When people say cluster they picture racks of GPUs. The GPUs are the easy part. What decides whether the cluster performs is everything around them:

  • Compute: N nodes, each 8 B300 SXM GPUs, with their own CPUs, system RAM and local NVMe.
  • Back-end fabric: the high-speed network that lets those nodes act as one machine.
  • Parallel storage: a filesystem fast enough to keep every GPU fed.
  • Front-end network and internet: how data and users reach the cluster.
  • Management and power: out-of-band control, and enough power and cooling to run the GPUs at full rated speed rather than capped.

Get any one of these wrong and you pay for GPUs that sit waiting.

The fabric decides everything

For multi-node training, the interconnect between nodes is the single most important choice after the GPUs themselves. During training the nodes constantly exchange gradients. If the fabric is slow or oversubscribed, the GPUs wait on the network instead of computing, and your expensive cluster scales badly. Ask:

  • InfiniBand (Quantum-X800) or Ethernet (Spectrum-X, RoCE v2)? Both can work, so ask which and why.
  • Non-blocking or oversubscribed? A non-blocking, rail-optimised fabric is what lets a job scale from 8 GPUs to thousands near-linearly. If it is oversubscribed, ask for the ratio.
  • What is the largest contiguous cluster on a single fabric?

We went deeper in the fabric that makes eight GPUs act as one. If a provider is vague about the fabric, that is a red flag on its own.

Storage that keeps the GPUs fed

A cluster full of B300s will stall if the storage cannot stream data fast enough. Ask for a parallel, GPU-aware filesystem (VAST, WEKA, DDN or similar), and ask for both capacity and aggregate throughput, not just terabytes. Throughput is what stops training pausing between batches.

How reserving allocation works, and how to do it safely

Blackwell GPUs are scarce and allocated months ahead, so any serious operator will ask you to commit to a term, usually with money down, before they hold a cluster for you. That is normal. Reserving allocation is not a scam. The scam is in how the money moves. A legitimate reservation looks like this:

  • Contract first, money second. A signed agreement that fixes the quantity, spec, term, price, a committed ready-for-service date, acceptance tests, and what happens if they miss it.
  • A proportionate deposit. A reservation fee or part-payment, often ten to thirty per cent or the first period prepaid, with the balance on delivery or on milestones. Not the whole lease wired up front.
  • Paid to a named company, against an invoice. A real legal entity and a company bank account, not an individual and not crypto to a wallet.
  • Escrow or milestones for large commitments, so your deposit is released as they deliver rather than gambled on trust.
  • Proof they control the capacity: a datacentre contract, an allocation letter, a delivery timeline.

The tell is not that they asked for a deposit. It is a deposit asked for before any contract, to an account nobody will stand behind, with no consequence if they fail to deliver. And if you are buying to resell to your own customer, never pay a supplier deposit until your customer has signed and paid theirs, ideally into escrow, so your money out is always covered by money in.

Be realistic about lead times

If someone promises a large B300 cluster in a few weeks, be sceptical. At cluster scale the hardware is supply-constrained, and provisioning, cabling and burn-in take real time. Weeks to months is normal; days is a warning sign. A phased delivery, a first block of nodes soon and the rest as capacity frees up, is often the honest answer, and it lets you start sooner.

What it costs, and why the term matters

Cluster pricing is driven by more than the GPUs. The fabric, the storage, the power and the length of your commitment all move the number. A longer reserved term buys a lower rate, because it gives the operator the demand certainty they built the cluster on. Weigh a reserved term against on-demand and against buying your own hardware in rent, buy or API, and watch the costs that hide outside the GPU-hour: storage, egress, setup and idle time. Our single-node rates are on the pricing page; a cluster is quoted to the workload.

Insist on acceptance testing

Never pay the balance on a cluster you have not tested. Before you accept:

  • NCCL all-reduce benchmarks across the full cluster, from the actual cluster you would get, not a datasheet, to prove the fabric scales.
  • A burn-in period to shake out the early hardware failures.
  • A defined SLA for uptime, and, just as important, how fast a failed GPU or node is replaced and whether they hold a spare pool.

Acceptance testing is your leverage. A real operator expects it.

Red flags when sourcing a cluster

  • A deposit demanded before any contract, or to a personal or overseas account.
  • Urgency and pressure: the allocation will be gone tomorrow, wire now.
  • A refusal to name the operator or the datacentre.
  • A timeline that is too good to be true.
  • No penalty or refund if they fail to deliver.

None of these on its own proves bad faith, but two or three together mean slow down.

How FlexiCloud fits

We would rather you asked all of this than not. Where we stand, plainly:

  • At present we do not own the datacentre, and we do not pretend to. We source dedicated B300 capacity from partners who own and operate the hardware in each region, stand up the cluster, and run it for you as bare metal with root access. You always know who the operator is.
  • On continuity, terms and deposits, we put the answers in writing before you commit. Ask us who holds your capacity, what protects a prepayment, and how you leave with your data.

If you are sizing a cluster, or not sure whether you need one, talk to an engineer. Compare the B300 and B200 nodes, or read why the right one lowers TCO in B300 vs B200. If a cluster is not the right answer for you, we will say so.

Written by The FlexiCloud Team

Frequently asked questions

Do I need a multi-node cluster, or will one B300 node do?
One 8-GPU B300 node pools a large amount of HBM, so many models train and serve on a single node. You only need multiple nodes when the model plus its KV cache will not fit in one node, or when training time on a single node is too long. If it fits, one node is cheaper and simpler.
InfiniBand or RoCE for a multi-node B300 cluster?
Both can work. What matters is that the back-end fabric is non-blocking and rail-optimised, so nodes exchange gradients at full speed and the cluster scales near-linearly. If the fabric is oversubscribed, ask for the ratio. A provider who is vague about the fabric is a warning sign.
How long does it take to stand up a B300 cluster?
At cluster scale, weeks to months, not days. Blackwell hardware is supply-constrained, and cabling and burn-in take real time. A promise of a large cluster in a few weeks is a red flag. Phased delivery, a first block of nodes soon and the rest later, is often the honest answer.
Is a deposit to reserve GPU allocation normal?
Yes. Scarce GPUs are allocated months ahead, so operators ask you to commit to a term with money down. A safe reservation has a signed contract first, a proportionate deposit (often 10 to 30 per cent or the first period), payment to a named company against an invoice, and escrow or milestones for large amounts. A deposit demanded before any contract, to a personal account, with no refund if they fail to deliver, is the scam version.
Can I get bare-metal root access on a rented cluster?
On a properly delivered cluster, yes. Ask whether it is bare metal or virtualised, and whether you control the OS, drivers, CUDA versions and orchestration. FlexiCloud delivers dedicated B300 nodes as bare metal with root access.

Want this handled for you?

We run the servers so you do not have to read the next one of these.