AI flagship · Hopper
NVIDIA H200 servers
Hopper with 141 GB of HBM3e: 70B-class models on a single GPU.
- 141 GBHBM3e per GPU
- 4.8 TB/smemory bandwidth
- NVLink 4GPU interconnect
- 1–8 GPUsper server
Price per GPU · monthly
$2,279/mo
−30%Market median $3,259≈ $3.12 per GPU-hour
Below every on-demand list price we found
Server size
Configure a H200 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the H200.
The H200 is the Hopper GPU of the H100 with much more and faster memory: 141 GB of HBM3e at 4.8 TB/s. That is enough to serve a 70B model in FP8 on one GPU, with room for the KV cache. It keeps Hopper’s strong FP64 and mature CUDA software support.
- 141 GB at 4.8 TB/s
1.8× the memory and 1.4× the bandwidth of the H100.
- Mature Hopper software
FP8 Transformer Engine, TensorRT-LLM, every major framework.
- Strong FP64
Good for simulation as well as AI.
From 1 to 8 H200 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× H200 | 141 GB | 24 | 256 GB | 3.84 TB | $2,279/mo | Configure 1× H200 | |
| 2× H200NVLink 4 | 282 GB | 48 | 512 GB | 7.68 TB | $4,558/mo | Configure 2× H200 | |
| 4× H200NVLink 4 | 564 GB | 96 | 1,024 GB | 15.36 TB | $9,116/mo | Configure 4× H200 | |
| 8× H200NVLink 4 | 1,128 GB | 192 | 2,048 GB | 30.72 TB | $18,232/mo | Configure 8× H200 |
NVIDIA H200 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Hopper
- Form factor
- SXM module
- GPU memory
- 141 GB HBM3e
- Memory bandwidth
- 4.8 TB/s
- GPU-to-GPU link
- NVLink 4, 900 GB/s per GPU
- Host link
- PCIe Gen5
- Max power
- Up to 700 W, configurable
- Multi-Instance GPU
- Up to 7 instances of 18 GB
Tensor throughput (dense)
- FP8
- 1,979 TFLOPS
- FP16 / BF16
- 989 TFLOPS
- TF32
- 495 TFLOPS
- INT8
- 1,979 TOPS
- FP32
- 67 TFLOPS
- FP64 / FP64 Tensor
- 34 / 67 TFLOPS
Cores and media
- Tensor Cores
- 4th generation
- Video and JPEG decoders
- 7 NVDEC, 7 NVJPG
Server, per GPU
- vCPU
- 24
- RAM
- 256 GB DDR5
- NVMe
- 3.84 TB
- GPUs per server
- 1, 2, 4, 8
NVIDIA publishes sparse tensor figures; dense figures are half of them. Hopper has no FP4 tensor cores.
Which models fit on H200 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× |
| Llama 3.3 70B | 2× | 1× | 1× |
| gpt-oss-120b | 4× | 2× | 1× |
| DeepSeek-V4-Flash | 8× | 4× | 2× |
| Qwen3.5-397B-A17B | 8× | 4× | 2× |
| DeepSeek-R1 (671B) | — | 8× | 4× |
| Kimi K2.6 (1T) | — | — | 8× |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest H200 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
H200 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
H200 questions.
More in the full FAQ.
Does a 70B model fit on one H200?
In FP8 a 70B model needs about 84 GB by our rule of thumb, which fits in 141 GB with room for the KV cache.
H200 or H100?
Same compute, but the H200 has 141 GB instead of 80 GB and more bandwidth. It serves larger models on fewer GPUs.
How many H200 per server?
1, 2, 4 or 8, connected by NVLink 4 at 900 GB/s per GPU.
Build your H200 server.
Pick 1 to 8 GPUs and see the exact monthly price. $2,279 per GPU, the same at every size.
