Questions to ask a GPU provider before you sign

The GPU rush has produced a lot of providers, and a lot of them are new. Some own datacentres full of hardware; many are intermediaries reselling capacity that belongs to someone else. Neither is automatically wrong, but the difference changes what you are actually buying and what happens if things go sideways. Before you lock in a six or twelve month commitment, here are the questions worth asking any GPU provider, in the order they matter. If a provider answers them straight, that tells you a lot. If they dodge the first few, that tells you more.
If you have not yet decided whether to rent at all, weigh it against buying your own hardware or using a pay-per-token API in rent, buy or API. This checklist is for once you have decided to rent, and need to choose who to rent from.
1. Who are you actually buying from?
This is the question most buyers skip, and it is the one that matters most.
- Do you own and operate this hardware, or are you reselling capacity from another provider? If reselling, who is the operator?
- Who owns the GPUs: you, a financing partner, or a lessor? What happens to our capacity if that arrangement changes?
- Is the capacity already racked, powered and burned in, or is it on order? What is the firm delivery date?
There is nothing wrong with buying through an intermediary. Plenty of good providers work that way, as long as they are transparent about it. What is not fine is a provider who implies they own infrastructure they are quietly reselling.
2. The question almost nobody asks: what if the company you are paying goes out of business?
If you are paying an intermediary but a different company owns and runs the metal, you carry a risk that rarely makes it into the sales conversation. If that intermediary fails mid-term, three things are exposed:
- Your prepayment. Prepay a year up front and, if they go under, you are an unsecured creditor in their insolvency. That money is usually gone.
- Your capacity. Your contract is with the intermediary, not the operator. If the intermediary stops paying the operator, the operator can reclaim your GPUs, and your training run dies because of someone else's cash-flow problem.
- Your data. It is sitting in a datacentre you may not even be able to name, under a contract that just evaporated.
So ask:
- Can I contract directly with the operator, or step into their contract, if you cannot continue?
- Is my capacity reserved in my name with the operator, or only in yours?
- Are my prepayments protected, or can we bill over the term instead of a large upfront?
- On any exit, how do I get my data out, and within how long?
A provider that has thought about your continuity will have clean answers. One that has not will change the subject.
3. The hardware
- Is it HGX B300 (8-GPU nodes) or GB300 NVL72 racks?
- Which OEM built the servers: Supermicro, Dell, HPE, other?
- Per node: CPU model and core count, system RAM, local NVMe capacity?
- Bare metal or virtualised? Do we get root access and control over drivers and CUDA versions?
- Air or liquid cooled? Are the GPUs running at full rated power, or power-capped?
That last one matters more than it sounds: a power-capped GPU is a slower GPU you are paying full price for.
4. Networking and storage
- Inter-node fabric: InfiniBand (Quantum-X800) or Ethernet (Spectrum-X)? Bandwidth per GPU?
- Is the fabric rail-optimised and non-blocking? If not, what is the oversubscription ratio?
- What is the largest contiguous cluster we can get on a single fabric?
- Shared storage: which platform (VAST, WEKA, DDN, other), what capacity and throughput?
- Internet bandwidth and egress pricing?
For multi-node training, the fabric decides whether your GPUs compute or wait. We went into why in the fabric that makes eight GPUs act as one.
5. Proof
- Can you share NCCL all-reduce benchmarks and burn-in reports for the actual cluster we would get?
- Can we run our own benchmarks on a test node before signing?
- Can we do a site visit or a live video walkthrough of the racks?
Real operators produce these without drama. Ask for the numbers from the actual cluster you would get, not a datasheet.
6. The datacentre
- Exact city and facility name. Your own building or colocation, and with which operator?
- Tier rating, power redundancy (N+1 or 2N), and dedicated MW for this cluster?
- PUE and cooling design?
- Certifications: SOC 2, ISO 27001? For India workloads, data residency and MeitY empanelment?
- Network latency from the datacentre to Mumbai and Chennai, or wherever your users are?
- What export-control approvals cover B300 deployment in this location, and who holds them?
Export controls are not a formality for Blackwell-class GPUs. Ask who holds the approvals, especially for anything crossing a border.
7. Commercial and support
- Price per GPU-hour on-demand, and for a one-year and a longer reserved term? Any prepayment required?
- Uptime SLA and credits? Guaranteed time to replace a failed GPU or node? Do you hold a spare pool?
- Support hours and time zone, and is there 24/7 remote hands at the site?
The short version
If a provider dodges questions 1, 2, 5 or 6, treat that as your answer. Who owns it, what happens if you disappear, prove it works, and where exactly it lives: real operators answer those without flinching.
How FlexiCloud answers these
We would rather you ask us all of this than not. Where we stand, plainly:
- Who you are buying from: at present we do not own the datacentre, and we do not pretend to. We source dedicated capacity from partners who own the hardware and hold it in each region, and we deliver and operate it for you as bare metal. You always know who the operator is.
- The hardware: dedicated 8-GPU HGX B300 and B200 SXM nodes, bare metal, root access, your own OS, drivers, CUDA and orchestration.
- What we run: we keep the hardware healthy for the whole term, monitored, failures handled, with an engineer you can reach, not a ticket queue.
- Continuity: we would rather you were never stranded by us. Ask us the four continuity questions above, who holds your capacity, what happens to a prepayment, and how you leave with your data, and we will put the answers in writing before you commit.
If a B300 or B200 node is the right fit, compare the two on the pricing page, or read the details for the B300 and B200. If it is not the right fit, we will tell you. Either way, talk to an engineer and ask us every question on this list.