GPU cloud · NVIDIA H100 SXM
Dedicated NVIDIA H100 SXM nodes, in India, the EU and the US
Dedicated 8-GPU HGX H100 nodes on a one-year term, the proven Hopper workhorse for training, fine-tuning and inference, at a lower rate than Blackwell. FlexiCloud delivers the bare metal and keeps the hardware running; you bring the workload and run your own stack.
From USD 2.95 per GPU-hour. Minimum one node of 8 GPUs on a one-year commitment.
The numbers that matter
- H100 GPUs per node
- 8
- HBM3 memory per GPU
- 80 GB
- Regions: India, EU, US
- 3
- Per GPU-hour, from
- $2.95
What we run, and what you run
The node is yours. The hardware is our job.
Most GPU providers hand you a bare node and a support ticket queue. We keep the hardware running underneath: monitored around the clock, failures handled, and an engineer you can actually reach.
Full control of the stack
Root access to run your own OS image, CUDA version, container runtime and orchestration. No platform layer between you and the GPUs; the software is entirely yours.
Watched around the clock
GPU health, thermals, memory errors and interconnect are monitored continuously. A failing card is caught by us, not discovered by a crashed training run at 3am.
An engineer, not a queue
Support is a person who runs GPU infrastructure, in your timezone. When something is wrong you talk to someone who can fix it.
Data where you need it
Nodes in India for domestic data residency, and in the EU and US for teams and datasets that live there. Same service, same team, whichever region.
Yours alone
Dedicated bare-metal nodes. No shared tenancy, no noisy neighbours, no scheduler deciding when your job runs. The whole node is yours for the term.
Predictable cost
A fixed term at a fixed rate. No spot-market surprises, no burst pricing, no bill that doubles because a job ran long. You know the number for the whole year.
Talk to an engineer
Get a quote for your workload
Leave a name and a number or email. An engineer calls you back with pricing and availability for your region, usually within one working day.
Terms
How it is sold
This is dedicated infrastructure on a term, not an hourly spot instance. The terms are simple and they are the same for everyone.
Minimum order
One node: 8 × H100 SXM
The unit is a full HGX H100 node. Eight GPUs with NVLink between them, dedicated to you. Need more? Nodes are added in units of eight.
- 8 × NVIDIA H100 SXM GPUs
- 80 GB HBM3 per GPU, 640 GB per node
- NVLink interconnect across the node
- Dedicated bare metal, single tenant
Pricing and term
From $2.95 per GPU-hour
On a one-year minimum commitment, billed for the term. Final pricing depends on region, term length and how many nodes, and we confirm it in writing before you commit.
- From $2.95 per GPU-hour
- One-year minimum commitment
- Longer terms and multi-node priced on request
Availability
India, the EU and the US
Capacity is allocated per region on a first-committed basis. Tell us where the data lives and where the team sits, and we will confirm availability and lead time for that region before you commit to anything.
- India: for domestic data residency
- European Union: for EU-resident data and teams
- United States: for US-based workloads
- Lead time confirmed per region before order
How it works
From first conversation to a running node
- 01
Tell us the workload
Training, fine-tuning or inference; the framework; how many nodes; and which region. Fifteen minutes with an engineer, not a form that disappears into a CRM.
- 02
We confirm capacity and price
A firm quote for your region and term, with the lead time to a live node. No commitment until you have the number in writing.
- 03
We hand over the node
A dedicated bare-metal node, networked and ready to log into with root access. You bring your own OS image, drivers and orchestration; the hardware underneath stays ours to keep running.
- 04
You run. We keep it healthy.
Your team owns the OS, the stack and the workload. Ours watches the hardware, handles failures, and is a phone call away for the whole term.
Questions
What people ask before they commit
The H100 is NVIDIA’s previous-generation Hopper GPU, with 80 GB of memory per GPU. The B200 and B300 are the newer Blackwell generation, with 192 GB and 288 GB per GPU and higher throughput. The H100 is the proven, cost-effective workhorse with the widest software and framework support; choose Blackwell when you need the extra memory for very large models and long context, or the newest performance.
One node, which is 8 NVIDIA H100 SXM GPUs. We do not sell fewer than a full node, because the GPUs share NVLink and the node is the unit of allocation. Additional capacity is added in whole nodes.
One year. This is dedicated bare-metal infrastructure reserved for you, not a spot instance, and the term is what makes the rate possible. Longer terms are available and are priced lower.
Pricing starts at $2.95 per GPU-hour, the floor rate on our longest commitments; it is higher on shorter terms and falls the longer you reserve. The minimum commitment is one year. The exact figure depends on your term length, region and the number of nodes, and we confirm it in writing before you commit.
India, the European Union and the United States. Capacity is allocated per region, so we confirm availability and lead time for your chosen region before you order. Pick the region by where your data has to live and where your team works.
We deliver the node as dedicated bare metal with network and root access, and keep the hardware running: continuous monitoring of GPU health and interconnect, hardware failure handling, and an engineer you can reach directly for the whole term. The software stack on top - OS, drivers, CUDA, container runtime and orchestration - is yours to run.
Yes. Nodes are added in units of eight GPUs, and multi-node deployments are priced on request. Tell us the target size when we scope the workload and we will plan capacity for it.
No. We work with partners who own the hardware and hold the capacity in each region. FlexiCloud sources the node, delivers it to you as dedicated bare metal, and keeps the hardware running - monitored, failures handled and supported - for the whole term. The software you run on top is yours.
Need more memory, or the newest generation?
The H100 is the proven, lower-cost choice. If you need more memory per GPU or Blackwell-generation throughput, the B200 and B300 step up from here.
Go deeper
More on running GPUs
The maths and the trade-offs behind an 8-GPU node, cost, memory, rent vs buy, and the interconnect that makes it one machine.
Hopper vs Blackwell: NVIDIA’s GPU generations compared
NVIDIA’s Hopper (H100, H200) and Blackwell (B200, B300) generations compared: memory, architecture, interconnect, power and price, and a practical guide to which one actually fits your AI workload.
Read itWhat it costs to fine-tune a model on an 8-GPU B300 node
The cost of a B300 job is GPU-hours × the rate, nothing more. Here is the actual arithmetic, three worked examples from a short LoRA run to a two-week training job, and an honest note on when a node is the wrong tool.
Read itWill your model fit? A practical guide to GPU memory for AI
Whether your model fits is arithmetic, not a benchmark. The bytes-per-parameter maths for inference and training, the levers that shrink it (quantization, LoRA, checkpointing), and a back-of-envelope guide to size any job, mapped to what fits on a B300.
Read itRent GPUs, buy your own, or use an API? A guide for teams building on AI
Pay-per-token API, buy your own GPUs, or rent bare-metal nodes? The right answer is about utilisation and control, not hype. A fair look at all three, and a simple way to decide which your workload actually needs.
Read it
Tell us the workload and the region
Fifteen minutes with an engineer gets you a firm price, a lead time, and a straight answer on whether the H100, B200 or B300 is the right fit. No commitment until you have the number in writing.