AI flagship · Hopper
NVIDIA H100 servers
The proven Hopper workhorse for training and inference.
- 80 GBHBM3 per GPU
- 3.35 TB/smemory bandwidth
- NVLink 4GPU interconnect
- 1–8 GPUsper server
Price per GPU · monthly
$1,779/mo
−30%Market median $2,548≈ $2.44 per GPU-hour
Below every on-demand list price we found
Server size
Configure a H100 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the H100.
The H100 SXM is the GPU most AI teams know best: 80 GB of HBM3, FP8 Transformer Engine and NVLink 4. Frameworks, kernels and recipes are tuned for it. It is the reference choice when you want predictable software behaviour at a lower price than newer generations.
- Best-supported AI GPU
Every framework and serving engine is tuned for Hopper.
- FP8 Transformer Engine
Mixed FP8 training and inference.
- NVLink 4 · 900 GB/s
Fast multi-GPU training inside the server.
From 1 to 8 H100 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× H100 | 80 GB | 20 | 200 GB | 2 TB | $1,779/mo | Configure 1× H100 | |
| 2× H100NVLink 4 | 160 GB | 40 | 400 GB | 4 TB | $3,558/mo | Configure 2× H100 | |
| 4× H100NVLink 4 | 320 GB | 80 | 800 GB | 8 TB | $7,116/mo | Configure 4× H100 | |
| 8× H100NVLink 4 | 640 GB | 160 | 1,600 GB | 16 TB | $14,232/mo | Configure 8× H100 |
NVIDIA H100 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Hopper, TSMC 4N
- Form factor
- SXM5 module
- GPU memory
- 80 GB HBM3, 5,120-bit
- Memory bandwidth
- 3.35 TB/s
- GPU-to-GPU link
- NVLink 4, 900 GB/s per GPU
- Host link
- PCIe Gen5
- Max power
- Up to 700 W, configurable
- Multi-Instance GPU
- Up to 7 instances of 10 GB
Tensor throughput (dense)
- FP8
- 1,979 TFLOPS
- FP16 / BF16
- 989 TFLOPS
- TF32
- 495 TFLOPS
- INT8
- 1,979 TOPS
- FP32
- 67 TFLOPS
- FP64 / FP64 Tensor
- 34 / 67 TFLOPS
Cores and media
- Streaming multiprocessors
- 132
- CUDA cores
- 16,896
- Tensor Cores
- 528, 4th generation
- Video and JPEG decoders
- 7 NVDEC, 7 NVJPG
Server, per GPU
- vCPU
- 20
- RAM
- 200 GB DDR5
- NVMe
- 2 TB
- GPUs per server
- 1, 2, 4, 8
NVIDIA publishes sparse tensor figures; dense figures are half of them. Hopper has no FP4 tensor cores.
Which models fit on H100 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 2× | 1× | 1× |
| Llama 3.3 70B | 4× | 2× | 1× |
| gpt-oss-120b | 4× | 2× | 2× |
| DeepSeek-V4-Flash | — | 8× | 4× |
| Qwen3.5-397B-A17B | — | 8× | 4× |
| DeepSeek-R1 (671B) | — | — | 8× |
| Kimi K2.6 (1T) | — | — | — |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest H100 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
H100 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
H100 questions.
More in the full FAQ.
Is the H100 still worth it in 2026?
Yes for most fine-tuning and inference work: it is fast, fully supported and costs less than newer GPUs. Choose more memory (H200, MI355X) when models or contexts outgrow 80 GB.
Which H100 do you rent?
The SXM version with 80 GB of HBM3 and NVLink 4, not the PCIe card.
What does an 8× H100 server include?
640 GB of GPU memory, 160 vCPU, 1,600 GB of RAM and 16 TB of NVMe.
Build your H100 server.
Pick 1 to 8 GPUs and see the exact monthly price. $1,779 per GPU, the same at every size.
