New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

AI flagship · Blackwell

NVIDIA B200 servers

Blackwell for training and high-throughput inference, with 180 GB of HBM3e per GPU.

  • 180 GBHBM3e per GPU
  • 8 TB/smemory bandwidth
  • NVLink 5GPU interconnect
  • 1–8 GPUsper server

Price per GPU · monthly

$3,439/mo

−30%

Market median $4,920≈ $4.71 per GPU-hour

Below every on-demand list price we found

Server size

Configure a B200 server

USD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC

Why the B200.

The B200 is NVIDIA’s Blackwell data-center GPU. Compared with Hopper it brings FP4 and FP8 tensor cores, 8 TB/s of memory bandwidth and NVLink 5 at 1.8 TB/s per GPU. An 8-GPU server holds 1.4 TB of HBM3e, enough for full fine-tuning of 70B-class models or serving several large models at once.

  • 8 TB/s memory bandwidth

    Token generation is bandwidth-bound; the B200 has 1.7× the bandwidth of the H200.

  • NVLink 5 · 1.8 TB/s

    Twice the GPU-to-GPU bandwidth of Hopper for sharded training.

  • FP4 and FP8

    Native low-precision tensor cores for inference and training.

From 1 to 8 B200 GPUs.

Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.

NVIDIA B200 server sizes and monthly prices
Server GPU memory vCPU RAM NVMe Market median CryptGPU Action
1× B200 180 GB 28 288 GB 3.84 TB $4,920 $3,439/mo Configure 1× B200
2× B200NVLink 5 360 GB 56 576 GB 7.68 TB $9,840 $6,878/mo Configure 2× B200
4× B200NVLink 5 720 GB 112 1,152 GB 15.36 TB $19,680 $13,756/mo Configure 4× B200
8× B200NVLink 5 1,440 GB 224 2,304 GB 30.72 TB $39,360 $27,512/mo Configure 8× B200

NVIDIA B200 specifications.

Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.

GPU and memory

Architecture
Blackwell, TSMC 4NP
Form factor
SXM module (HGX B200)
GPU memory
180 GB HBM3e
Memory bandwidth
Up to 8 TB/s
GPU-to-GPU link
NVLink 5, 1.8 TB/s per GPU
Host link
PCIe Gen5
Max power
Up to 1,000 W, configurable
Multi-Instance GPU
Up to 7 instances

Tensor throughput (dense)

FP4
9 PFLOPS
FP8 / FP6
4.5 PFLOPS
FP16 / BF16
2.25 PFLOPS
TF32
1.1 PFLOPS
INT8
4.5 POPS
FP32
75 TFLOPS
FP64
37 TFLOPS

Cores and media

Tensor Cores
5th generation
Video decoders
7 NVDEC

Server, per GPU

vCPU
28
RAM
288 GB DDR5
NVMe
3.84 TB
GPUs per server
1, 2, 4, 8

NVIDIA publishes sparse tensor figures; dense figures are half of them. Memory bandwidth is 8 TB/s on NVIDIA’s DGX B200 page and 7.7 TB/s in its Blackwell datasheet.

Which models fit on B200 servers.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.

Number of B200 GPUs needed per model and precision
Model (total parameters)FP16 / BF16FP84-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest B200 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.

B200 prices across the market.

Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.

  • CryptGPUmonthly, dedicated$3,439$4.71/h
  • Lowest list priceon-demand, 10 providers$3,964$5.43/h
  • Market medianon-demand$4,920$6.74/h
  • Highest list priceon-demand$6,278$8.60/h
  • Oracle Cloudlist price · reference$10,220$14.00/h
  • AWSp6-b200.48xlarge · reference$10,395$14.24/h

Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare

B200 questions.

More in the full FAQ.

B200 or H200?

The B200 has more memory (180 GB vs 141 GB), 1.7× the bandwidth and twice the NVLink bandwidth, plus FP4. The H200 costs less per GPU and has stronger FP64.

How much GPU memory does an 8× B200 server have?

1,440 GB of HBM3e, with 224 vCPU, 2,304 GB of RAM and 30.72 TB of NVMe.

Is the B200 a good inference GPU?

Yes. One B200 holds a 70B model in FP8 (about 84 GB by our rule of thumb) with plenty of memory left for the KV cache, and 8 TB/s of bandwidth speeds up token generation. Engines that support FP4 on Blackwell cut memory needs further.

Build your B200 server.

Pick 1 to 8 GPUs and see the exact monthly price. $3,439 per GPU, the same at every size.