GPU comparison · Hopper vs Hopper
NVIDIA H100 vs H200.
Same Hopper compute, different memory. The H200 carries 141 GB of HBM3e at 4.8 TB/s against 80 GB at 3.35 TB/s on the H100: it holds 70B-class models on one GPU and feeds the cores faster, for $500 more per GPU per month.
- 80 vs 141 GBGPU memory
- 3.35 vs 4.8 TB/smemory bandwidth
- $1,779 vs $2,279per GPU, per month
H100 vs H200: which one to rent.
Choose the H100 if
- Your model, its KV cache and batch fit in 80 GB per GPU (up to about 60B parameters in FP8 by our sizing rule).
- Your jobs are compute-bound, such as training or fine-tuning mid-size models: both GPUs have the same tensor throughput.
- You want the lowest monthly price per Hopper GPU: $1,779 against $2,279.
Choose the H200 if
- You serve 70B-class models: about 84 GB in FP8, which fits on one H200 but needs two H100.
- You run long contexts or large batches: 61 GB more per GPU goes to the KV cache.
- Token generation is memory-bound in your workload: 43% more bandwidth feeds the same cores faster.
H100 vs H200 specs, side by side.
Figures from NVIDIA’s published specifications, per GPU. Tensor figures are peak dense throughput.
| Per GPU | NVIDIA H100 | NVIDIA H200 | H200 / H100 |
|---|---|---|---|
| Architecture | Hopper, TSMC 4N | Hopper | — |
| Form factor | SXM5 module | SXM module | — |
| GPU memory | 80 GB HBM3, 5,120-bit | 141 GB HBM3e | 1.76× |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s | 1.43× |
| GPU-to-GPU link | NVLink 4, 900 GB/s per GPU | NVLink 4, 900 GB/s per GPU | Same |
| FP4 tensor (dense) | Not supported | Not supported | — |
| FP8 tensor (dense) | 1,979 TFLOPS | 1,979 TFLOPS | Same |
| FP16 / BF16 tensor (dense) | 989 TFLOPS | 989 TFLOPS | Same |
| FP64 | 34 / 67 TFLOPS | 34 / 67 TFLOPS | Same |
| Max power | Up to 700 W, configurable | Up to 700 W, configurable | — |
| GPUs per server | 1, 2, 4, 8 | 1, 2, 4, 8 | — |
| Per GPU in our servers | 20 vCPU · 200 GB RAM · 2 TB | 24 vCPU · 256 GB RAM · 3.84 TB | — |
Sources: NVIDIA H100 · NVIDIA H200. Full sheets: H100 · H200
H100 vs H200 rental price.
Our monthly prices, set 30% below the market median of public on-demand prices and paid in crypto. Hourly figures are the monthly price divided by 730 hours.
| CryptGPU | NVIDIA H100 | NVIDIA H200 | H200 / H100 |
|---|---|---|---|
| Price per GPU, per month | $1,779 | $2,279 | 1.28× |
| Equivalent per GPU-hour | $2.44 | $3.12 | — |
| Per GB of GPU memory, per month | $22.24 | $16.16 | 0.73× |
| Per PFLOPS of dense FP8, per month | $899 | $1,152 | 1.28× |
| Largest server | 8× · $14,232/mo | 8× · $18,232/mo | — |
| Market median, per GPU | $2,548 | $3,259 | — |
| Below the median | −30% | −30% | — |
Same price per GPU at every server size, no hourly metering. How we compare prices · Configure an H100 server · Configure a H200 server
Which models fit on each.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache, 92% of GPU memory usable.
| Model (total parameters) | H100 · FP8 | H200 · FP8 | H100 · 4-bit | H200 · 4-bit |
|---|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× | 1× |
| Llama 3.3 70B | 2× | 1× | 1× | 1× |
| gpt-oss-120b | 2× | 2× | 2× | 1× |
| DeepSeek-V4-Flash | 8× | 4× | 4× | 2× |
| Qwen3.5-397B-A17B | 8× | 4× | 4× | 2× |
| DeepSeek-R1 (671B) | — | 8× | 8× | 4× |
| Kimi K2.6 (1T) | — | — | — | 8× |
FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter. — = larger than the biggest server of that GPU. Size another model · How much VRAM does an LLM need?
H100 vs H200 FAQ.
More on each GPU: NVIDIA H100 · NVIDIA H200.
Is the H200 faster than the H100?
The compute is the same Hopper die, so compute-bound jobs run at a similar speed. The H200 is faster where memory limits the work, as in LLM token generation, thanks to 4.8 TB/s against 3.35 TB/s, and it can keep larger models and batches on fewer GPUs.
How much more does an H200 cost to rent?
Per GPU per month, $2,279 against $1,779. Per GB of GPU memory the H200 is cheaper: $16.16 against $22.24 a month.
Can a 70B model run on one H100?
Not in FP8: it needs about 84 GB by our rule of thumb, more than 80 GB. In 4-bit it needs about 46 GB and fits on one H100. In FP8 it fits on one H200 or on two H100.
Do the H100 and H200 use the same software?
Yes. Both are Hopper GPUs with the FP8 Transformer Engine, NVLink 4 and the same CUDA, cuDNN, TensorRT-LLM, vLLM and SGLang support.
Other GPU comparisons.
Rent the H100 or the H200.
Dedicated servers with 1 to 8 GPUs, one monthly price, paid in crypto. Online in under 10 minutes, no KYC.
