NVIDIA · Blackwell Ultra
B300
New288 GB HBM3e
- Bandwidth
- 8 TB/s
- Link
- NVLink 5
- Server
- 1–8 GPUs
Frontier training, 400B+ inference, long context
New NVIDIA B300 — from $4,019/mo
See the B300Menu
Pay with BTCETHUSDTUSDCLTC
10 GPU models · 1 to 8 GPUs per server
From the GeForce RTX 5080 to the NVIDIA B300 and the AMD MI355X. Every server is dedicated to you, with root access and one flat monthly price about 30% under the market median.
Filter by family. Each card opens the full specifications, every server size and the market check for that GPU.
NVIDIA · Blackwell Ultra
288 GB HBM3e
Frontier training, 400B+ inference, long context
NVIDIA · Blackwell
180 GB HBM3e
Large-scale training and high-throughput inference
AMD · CDNA 4
288 GB HBM3E
Memory-hungry inference, ROCm training, FP64 science
NVIDIA · Hopper
141 GB HBM3e
70B-class models on one GPU, fine-tuning, serving
NVIDIA · Hopper
80 GB HBM3
The proven workhorse for training and inference
NVIDIA · Blackwell
96 GB GDDR7 ECC
96 GB per card for inference, fine-tuning, rendering
NVIDIA · Ada Lovelace
48 GB GDDR6 ECC
Cost-efficient inference, diffusion, video pipelines
NVIDIA · Blackwell
32 GB GDDR7
Image & video generation, 3D rendering, dev boxes
NVIDIA · Ada Lovelace
24 GB GDDR6X
Budget inference, Stable Diffusion, rendering
NVIDIA · Blackwell
16 GB GDDR7
Entry GPU for small models, CI, development, video encoding
Sort by price, memory, bandwidth or price per GB of GPU memory. Prices are per GPU per month; a server of n GPUs costs n times as much.
| Interconnect | Precisions | GPUs / server | Market median | |||||
|---|---|---|---|---|---|---|---|---|
| NVIDIA B300Blackwell Ultra | 288 GBHBM3e | 8 TB/s | NVLink 5 · 1.8 TB/s | FP4 · FP8 · BF16 | 1–8 | $4,019 | $13.95 | |
| NVIDIA B200Blackwell | 180 GBHBM3e | 8 TB/s | NVLink 5 · 1.8 TB/s | FP4 · FP8 · BF16 | 1–8 | $3,439 | $19.11 | |
| AMD MI355XCDNA 4 | 288 GBHBM3E | 8 TB/s | Infinity Fabric · 1.07 TB/s | FP4 · FP6 · FP8 | 1–8 | $1,409 | $4.89 | |
| NVIDIA H200Hopper | 141 GBHBM3e | 4.8 TB/s | NVLink 4 · 900 GB/s | FP8 · BF16 · FP64 | 1–8 | $2,279 | $16.16 | |
| NVIDIA H100Hopper | 80 GBHBM3 | 3.35 TB/s | NVLink 4 · 900 GB/s | FP8 · BF16 · FP64 | 1–8 | $1,779 | $22.24 | |
| NVIDIA RTX PRO 6000Blackwell | 96 GBGDDR7 ECC | 1.6 TB/s | PCIe Gen5 x16 | FP4 · FP8 · BF16 | 1–8 | $959 | $9.99 | |
| NVIDIA L40SAda Lovelace | 48 GBGDDR6 ECC | 864 GB/s | PCIe Gen4 x16 | FP8 · BF16 | 1–8 | $789 | $16.44 | |
| NVIDIA RTX 5090Blackwell | 32 GBGDDR7 | 1.8 TB/s | PCIe Gen5 x16 | FP4 · FP8 · BF16 | 1–4 | $329 | $10.28 | |
| NVIDIA RTX 4090Ada Lovelace | 24 GBGDDR6X | 1 TB/s | PCIe Gen4 x16 | FP8 · BF16 | 1–4 | $269 | $11.21 | |
| NVIDIA RTX 5080Blackwell | 16 GBGDDR7 | 960 GB/s | PCIe Gen5 x16 | FP4 · FP8 · BF16 | 1–4 | $199 | $12.44 |
$ / GB = monthly price per GB of GPU memory. Market median = median on-demand list price per GPU from 40 GPU clouds, 23 Sep 2026 × 730 hours. Full market comparison
Three questions settle most choices. The matrix shows where each GPU fits best.
Weights, KV cache and activations must fit in GPU memory. 16 GB on the RTX 5080, up to 288 GB per GPU on the B300 and MI355X, and 2.3 TB in an 8-GPU server. Size a model
Token and image generation are mostly limited by memory bandwidth: up to 8 TB/s on HBM flagships, 1.6 to 1.8 TB/s on the RTX PRO 6000 and RTX 5090, under 1 TB/s on the L40S and RTX 5080.
Sharded training exchanges data between GPUs at every step: NVLink (900 GB/s–1.8 TB/s) and Infinity Fabric are built for it. PCIe cards suit independent jobs and inference.
| GPU | LLM training | Fine-tuning | Inference & serving | Image & video | 3D rendering & VFX | Research & HPC |
|---|---|---|---|---|---|---|
| B300 | Recommended | Good fit | Good fit | — | — | — |
| B200 | Recommended | Good fit | Good fit | — | — | — |
| MI355X | Recommended | — | Recommended | — | — | Recommended |
| H200 | — | Recommended | Recommended | — | — | Recommended |
| H100 | Good fit | Recommended | Good fit | — | — | Recommended |
| RTX PRO 6000 | — | Recommended | Good fit | Recommended | Recommended | — |
| L40S | — | — | Recommended | Recommended | Good fit | — |
| RTX 5090 | — | — | Good fit | Recommended | Recommended | — |
| RTX 4090 | — | — | Good fit | Good fit | Recommended | — |
| RTX 5080 | — | — | Good fit | Good fit | — | — |
RecommendedGood fit
Billing and payments are covered in the full FAQ.
Start from memory: the model, its context and the batch must fit in GPU memory. Then look at memory bandwidth, which sets how fast tokens or images are generated, and at the interconnect if you train on several GPUs. The sizing helper lists every configuration that fits a given model, cheapest first.
The B300, B200, MI355X, H200 and H100 use stacked HBM memory (3.35 to 8 TB/s) and a high-speed link between the GPUs of a server: NVLink or Infinity Fabric. The RTX PRO 6000, L40S and GeForce cards use GDDR memory and talk to each other over PCIe. HBM platforms are faster for training and very large models; PCIe cards cost much less for serving, generation and rendering.
They are fast and cost-efficient for image and video generation, rendering, development and smaller models. They have no ECC memory and no NVLink, and servers take up to four cards. For long-running services that need ECC memory, choose the RTX PRO 6000 or the L40S.
Yes: order the new configuration, move your data, and let the old server’s term end. Each server keeps its own monthly term; moving data between servers is your responsibility.
The range focuses on current generations: Blackwell, Blackwell Ultra, Hopper, CDNA 4 and the latest GeForce and RTX PRO cards. The L40S and RTX 4090 (Ada Lovelace) stay in the range as budget options.
Pick 1 to 8 GPUs and see the exact monthly price, next to the market median.