AI flagship · Blackwell
NVIDIA B200 servers
Blackwell for training and high-throughput inference, with 180 GB of HBM3e per GPU.
- 180 GBHBM3e per GPU
- 8 TB/smemory bandwidth
- NVLink 5GPU interconnect
- 1–8 GPUsper server
Price per GPU · monthly
$3,439/mo
−30%Market median $4,920≈ $4.71 per GPU-hour
Below every on-demand list price we found
Server size
Configure a B200 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the B200.
The B200 is NVIDIA’s Blackwell data-center GPU. Compared with Hopper it brings FP4 and FP8 tensor cores, 8 TB/s of memory bandwidth and NVLink 5 at 1.8 TB/s per GPU. An 8-GPU server holds 1.4 TB of HBM3e, enough for full fine-tuning of 70B-class models or serving several large models at once.
- 8 TB/s memory bandwidth
Token generation is bandwidth-bound; the B200 has 1.7× the bandwidth of the H200.
- NVLink 5 · 1.8 TB/s
Twice the GPU-to-GPU bandwidth of Hopper for sharded training.
- FP4 and FP8
Native low-precision tensor cores for inference and training.
From 1 to 8 B200 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× B200 | 180 GB | 28 | 288 GB | 3.84 TB | $3,439/mo | Configure 1× B200 | |
| 2× B200NVLink 5 | 360 GB | 56 | 576 GB | 7.68 TB | $6,878/mo | Configure 2× B200 | |
| 4× B200NVLink 5 | 720 GB | 112 | 1,152 GB | 15.36 TB | $13,756/mo | Configure 4× B200 | |
| 8× B200NVLink 5 | 1,440 GB | 224 | 2,304 GB | 30.72 TB | $27,512/mo | Configure 8× B200 |
NVIDIA B200 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Blackwell, TSMC 4NP
- Form factor
- SXM module (HGX B200)
- GPU memory
- 180 GB HBM3e
- Memory bandwidth
- Up to 8 TB/s
- GPU-to-GPU link
- NVLink 5, 1.8 TB/s per GPU
- Host link
- PCIe Gen5
- Max power
- Up to 1,000 W, configurable
- Multi-Instance GPU
- Up to 7 instances
Tensor throughput (dense)
- FP4
- 9 PFLOPS
- FP8 / FP6
- 4.5 PFLOPS
- FP16 / BF16
- 2.25 PFLOPS
- TF32
- 1.1 PFLOPS
- INT8
- 4.5 POPS
- FP32
- 75 TFLOPS
- FP64
- 37 TFLOPS
Cores and media
- Tensor Cores
- 5th generation
- Video decoders
- 7 NVDEC
Server, per GPU
- vCPU
- 28
- RAM
- 288 GB DDR5
- NVMe
- 3.84 TB
- GPUs per server
- 1, 2, 4, 8
NVIDIA publishes sparse tensor figures; dense figures are half of them. Memory bandwidth is 8 TB/s on NVIDIA’s DGX B200 page and 7.7 TB/s in its Blackwell datasheet.
Which models fit on B200 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× |
| Llama 3.3 70B | 2× | 1× | 1× |
| gpt-oss-120b | 2× | 1× | 1× |
| DeepSeek-V4-Flash | 8× | 4× | 2× |
| Qwen3.5-397B-A17B | 8× | 4× | 2× |
| DeepSeek-R1 (671B) | — | 8× | 4× |
| Kimi K2.6 (1T) | — | 8× | 4× |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest B200 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
B200 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
B200 questions.
More in the full FAQ.
B200 or H200?
The B200 has more memory (180 GB vs 141 GB), 1.7× the bandwidth and twice the NVLink bandwidth, plus FP4. The H200 costs less per GPU and has stronger FP64.
How much GPU memory does an 8× B200 server have?
1,440 GB of HBM3e, with 224 vCPU, 2,304 GB of RAM and 30.72 TB of NVMe.
Is the B200 a good inference GPU?
Yes. One B200 holds a 70B model in FP8 (about 84 GB by our rule of thumb) with plenty of memory left for the KV cache, and 8 TB/s of bandwidth speeds up token generation. Engines that support FP4 on Blackwell cut memory needs further.
Build your B200 server.
Pick 1 to 8 GPUs and see the exact monthly price. $3,439 per GPU, the same at every size.
