New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Pro & inference · Blackwell

NVIDIA RTX PRO 6000 servers

Blackwell with 96 GB of ECC GDDR7 per card, for serving, fine-tuning and rendering.

  • 96 GBGDDR7 ECC per GPU
  • 1.6 TB/smemory bandwidth
  • PCIe Gen5 x16GPU interconnect
  • 1–8 GPUsper server

Price per GPU · monthly

$959/mo

−31%

Market median $1,380≈ $1.31 per GPU-hour

Below every on-demand list price we found

Server size

Configure a RTX PRO 6000 server

USD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC

Why the RTX PRO 6000.

The RTX PRO 6000 Blackwell is NVIDIA’s professional GPU for data centers and workstations. With 96 GB of ECC GDDR7 per card, FP4/FP8 tensor cores and RT cores, it covers LLM serving, QLoRA fine-tuning, diffusion, video and rendering at a much lower price than HBM GPUs.

  • 96 GB of ECC memory

    More memory per card than the H100, at about half the price.

  • Blackwell FP4 and FP8

    Modern low-precision inference.

  • RT cores and video engines

    Rendering and video pipelines as well as AI.

From 1 to 8 RTX PRO 6000 GPUs.

Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.

NVIDIA RTX PRO 6000 server sizes and monthly prices
Server GPU memory vCPU RAM NVMe Market median CryptGPU Action
1× RTX PRO 6000 96 GB 24 180 GB 1.92 TB $1,380 $959/mo Configure 1× RTX PRO 6000
2× RTX PRO 6000PCIe Gen5 x16 192 GB 48 360 GB 3.84 TB $2,760 $1,918/mo Configure 2× RTX PRO 6000
4× RTX PRO 6000PCIe Gen5 x16 384 GB 96 720 GB 7.68 TB $5,520 $3,836/mo Configure 4× RTX PRO 6000
8× RTX PRO 6000PCIe Gen5 x16 768 GB 192 1,440 GB 15.36 TB $11,040 $7,672/mo Configure 8× RTX PRO 6000

NVIDIA RTX PRO 6000 specifications.

Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.

GPU and memory

Architecture
Blackwell
Edition
Server Edition, passive PCIe card
GPU memory
96 GB GDDR7 with ECC, 512-bit
Memory bandwidth
1,597 GB/s
GPU-to-GPU link
None (PCIe 5.0 x16)
Max power
Up to 600 W, configurable
Multi-Instance GPU
Up to 4 instances of 24 GB

Tensor throughput (published)

FP4
4 PFLOPS
FP8
2 PFLOPS
FP16 / BF16
1 PFLOPS
TF32
234 TFLOPS
FP32
120 TFLOPS

Cores and media

CUDA cores
24,064
Tensor Cores
752, 5th generation
RT Cores
188, 4th generation
Video engines
4 NVENC, 4 NVDEC, 4 NVJPG

Server, per GPU

vCPU
24
RAM
180 GB DDR5
NVMe
1.92 TB
GPUs per server
1, 2, 4, 8

Figures for the Server Edition. NVIDIA does not say whether its tensor figures for this edition use sparsity; the equivalent Workstation Edition figures do.

Which models fit on RTX PRO 6000 servers.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.

Number of RTX PRO 6000 GPUs needed per model and precision
Model (total parameters)FP16 / BF16FP84-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX PRO 6000 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.

RTX PRO 6000 prices across the market.

Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.

  • CryptGPUmonthly, dedicated$959$1.31/h
  • Lowest list priceon-demand, 9 providers$975$1.34/h
  • Market medianon-demand$1,380$1.89/h
  • Highest list priceon-demand$2,190$3.00/h
  • AWSg7e.2xlarge · reference$2,455$3.36/h
  • Google Cloudg4-standard-48 · reference$3,285$4.50/h
  • Oracle Cloudlist price · reference$3,285$4.50/h

Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare

RTX PRO 6000 questions.

More in the full FAQ.

RTX PRO 6000 or H100?

The RTX PRO 6000 has more memory (96 GB vs 80 GB) and costs about half as much; the H100 has far more bandwidth, NVLink and faster training. For inference and fine-tuning the RTX PRO 6000 is often the better value.

Can it fine-tune a 70B model?

With QLoRA, yes: about 53 GB by our rule of thumb, well within 96 GB.

How many per server?

1, 2, 4 or 8 cards.

Build your RTX PRO 6000 server.

Pick 1 to 8 GPUs and see the exact monthly price. $959 per GPU, the same at every size.