Skip to content

GPU cloud · NVIDIA H200 SXM

Dedicated NVIDIA H200 SXM nodes, in India, the EU and the US

Dedicated 8-GPU HGX H200 nodes on a one-year term, the memory-upgraded Hopper with 141 GB of HBM3e per GPU, for inference, long context and larger models at a step below Blackwell. FlexiCloud delivers the bare metal and keeps the hardware running; you bring the workload and run your own stack.

From USD 3.35 per GPU-hour. Minimum one node of 8 GPUs on a one-year commitment.

The numbers that matter

H200 GPUs per node
8
HBM3e memory per GPU
141 GB
Regions: India, EU, US
3
Per GPU-hour, from
$3.35

Built for

Where the H200 earns its place

The H200 is the memory-upgraded Hopper: the same compute as the H100 with nearly double the memory and bandwidth, for workloads the H100 runs out of room for.

Large-model inference

Serve bigger models on fewer GPUs

141 GB of HBM3e holds larger models and longer context than the H100’s 80 GB, so you serve them without sharding across extra cards, on the same CUDA stack you already run.

  • Large-model and long-context serving
  • Fewer GPUs per model than the H100
  • Same drivers and CUDA as the H100

Memory-bound workloads

More headroom at Hopper cost

The H200 matches the H100 on compute but adds roughly 1.4x the bandwidth (around 4.8 TB/s) and nearly double the memory, which lifts throughput wherever the bottleneck is memory rather than maths.

  • Memory-bound fine-tuning
  • Higher inference throughput than the H100
  • About 4.8 TB/s of memory bandwidth

When to pick it

Choose the H200 when the H100 runs out of memory

Reach for the H200 when your model or context will not fit in 80 GB but you do not need Blackwell’s price or FP4. If 80 GB is enough, the H100 is cheaper; if you need more still, the B200 steps up.

  • 141 GB HBM3e per GPU
  • A step below Blackwell pricing
  • Same Hopper generation as the H100

What we run, and what you run

The node is yours. The hardware is our job.

Most GPU providers hand you a bare node and a support ticket queue. We keep the hardware running underneath: monitored around the clock, failures handled, and an engineer you can actually reach.

  • Full control of the stack

    Root access to run your own OS image, CUDA version, container runtime and orchestration. No platform layer between you and the GPUs; the software is entirely yours.

  • Memory where it counts

    141 GB of HBM3e per GPU, nearly double the H100, at roughly 4.8 TB/s. Larger models and longer context fit without sharding, and memory-bound inference runs faster.

  • Watched around the clock

    GPU health, thermals, memory errors and interconnect are monitored continuously. A failing card is caught by us, not discovered by a crashed training run at 3am.

  • An engineer, not a queue

    Support is a person who runs GPU infrastructure, in your timezone. When something is wrong you talk to someone who can fix it.

  • Data where you need it

    Nodes in India for domestic data residency, and in the EU and US for teams and datasets that live there. Same service, same team, whichever region.

  • Predictable cost

    A fixed term at a fixed rate. No spot-market surprises, no burst pricing, no bill that doubles because a job ran long. You know the number for the whole year.

Talk to an engineer

Get a quote for your workload

Leave a name and a number or email. An engineer calls you back with pricing and availability for your region, usually within one working day.

Terms

How it is sold

This is dedicated infrastructure on a term, not an hourly spot instance. The terms are simple and they are the same for everyone.

Minimum order

One node: 8 × H200 SXM

The unit is a full HGX H200 node. Eight GPUs with NVLink between them, dedicated to you. Need more? Nodes are added in units of eight.

  • 8 × NVIDIA H200 SXM GPUs
  • 141 GB HBM3e per GPU, 1,128 GB per node
  • NVLink interconnect across the node
  • Dedicated bare metal, single tenant
One-year minimum

Pricing and term

From $3.35 per GPU-hour

On a one-year minimum commitment, billed for the term. Final pricing depends on region, term length and how many nodes, and we confirm it in writing before you commit.

  • From $3.35 per GPU-hour
  • One-year minimum commitment
  • Longer terms and multi-node priced on request

Availability

India, the EU and the US

Capacity is allocated per region on a first-committed basis. Tell us where the data lives and where the team sits, and we will confirm availability and lead time for that region before you commit to anything.

  • India: for domestic data residency
  • European Union: for EU-resident data and teams
  • United States: for US-based workloads
  • Lead time confirmed per region before order

How it works

From first conversation to a running node

  1. 01

    Tell us the workload

    Training, fine-tuning or inference; the framework; how many nodes; and which region. Fifteen minutes with an engineer, not a form that disappears into a CRM.

  2. 02

    We confirm capacity and price

    A firm quote for your region and term, with the lead time to a live node. No commitment until you have the number in writing.

  3. 03

    We hand over the node

    A dedicated bare-metal node, networked and ready to log into with root access. You bring your own OS image, drivers and orchestration; the hardware underneath stays ours to keep running.

  4. 04

    You run. We keep it healthy.

    Your team owns the OS, the stack and the workload. Ours watches the hardware, handles failures, and is a phone call away for the whole term.

Questions

What people ask before they commit

  • The H200 is the same Hopper generation and the same compute as the H100, with one big change: memory. It has 141 GB of HBM3e per GPU at roughly 4.8 TB/s, against the H100’s 80 GB of HBM3 at around 3.4 TB/s. That nearly doubles the room per GPU, so larger models and longer context fit without sharding, and memory-bound inference runs noticeably faster. If your workload is limited by memory rather than raw compute, the H200 is the upgrade, for a small step up in price.

  • The H200 is Hopper; the B200 and B300 are the newer Blackwell generation, with 192 GB and 288 GB per GPU, a 4-bit inference format and roughly double the NVLink bandwidth. The H200 is the value step when you need more memory than an H100 gives you but not Blackwell’s price or performance. Choose Blackwell for the very largest models, the heaviest training, or when FP4 inference changes your economics.

  • One node, which is 8 NVIDIA H200 SXM GPUs. We do not sell fewer than a full node, because the GPUs share NVLink and the node is the unit of allocation. Additional capacity is added in whole nodes.

  • One year. This is dedicated bare-metal infrastructure reserved for you, not a spot instance, and the term is what makes the $3.35 per GPU-hour floor possible. Longer terms are priced lower.

  • Pricing starts at $3.35 per GPU-hour, the floor rate on our longest commitments; it is higher on shorter terms and falls the longer you reserve. The minimum commitment is one year. The exact figure depends on your term length, region and the number of nodes, and we confirm it in writing before you commit.

  • The H200 is NVIDIA’s Hopper generation, refreshed. It has the same compute as the H100 but 141 GB of HBM3e per GPU at roughly 4.8 TB/s, against the H100’s 80 GB at around 3.4 TB/s. It uses the Transformer Engine with FP8, and fourth-generation NVLink at roughly 900 GB/s per GPU. An 8-GPU HGX H200 node holds 1,128 GB of HBM3e.

  • India, the European Union and the United States. The H200 is in steadier supply than the Blackwell parts, but capacity is still allocated per region, so we confirm H200 availability and lead time for your region before you order. Pick the region by where your data has to live and where your team works.

  • We deliver the node as dedicated bare metal with network and root access, and keep the hardware running: continuous monitoring of GPU health and interconnect, hardware failure handling, and an engineer you can reach directly for the whole term. The software stack on top - OS, drivers, CUDA, container runtime and orchestration - is yours to run.

  • Yes. Nodes are added in units of eight H200 GPUs (1,128 GB of HBM3e each), and multi-node clusters are wired with a dedicated InfiniBand fabric, priced on request. Tell us the target size when we scope the workload and we will plan the interconnect and capacity for it.

  • No. We work with partners who own the hardware and hold the capacity in each region. FlexiCloud sources the node, delivers it to you as dedicated bare metal, and keeps the hardware running - monitored, failures handled and supported - for the whole term. The software you run on top is yours.

Need it cheaper, or need Blackwell?

The H200 is the memory-upgraded Hopper. If cost matters more than memory, the H100 is lower; if you need more memory still or Blackwell-generation throughput, the B200 and B300 step up from here.

Tell us the workload and the region

Fifteen minutes with an engineer gets you a firm price, a lead time, and a straight answer on whether the H100, H200, B200 or B300 is the right fit. No commitment until you have the number in writing.