Creator & consumer · Blackwell
NVIDIA RTX 5090 servers Hot
The fastest GeForce card: 32 GB of GDDR7 for generation, rendering and development.
- 32 GBGDDR7 per GPU
- 1.8 TB/smemory bandwidth
- PCIe Gen5 x16GPU interconnect
- 1–4 GPUsper server
Price per GPU · monthly
$329/mo
−32%Market median $481≈ $0.45 per GPU-hour
Below every on-demand list price we found
Server size
Configure a RTX 5090 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the RTX 5090.
The GeForce RTX 5090 is NVIDIA’s top consumer GPU of the Blackwell generation. Its 32 GB of GDDR7 at 1.8 TB/s and FP4/FP8 tensor cores make it the best price per image for diffusion models, a strong renderer and an affordable machine for model development.
- 1.8 TB/s GDDR7
More memory bandwidth than the RTX PRO 6000 Server Edition (1.6 TB/s), at about a third of the price.
- 32 GB
Enough for most image models and 30B-class LLMs in 4-bit.
- Up to 4 per server
Run parallel workers or multi-GPU renders.
From 1 to 4 RTX 5090 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× RTX 5090 | 32 GB | 16 | 64 GB | 1 TB | $329/mo | Configure 1× RTX 5090 | |
| 2× RTX 5090PCIe Gen5 x16 | 64 GB | 32 | 128 GB | 2 TB | $658/mo | Configure 2× RTX 5090 | |
| 4× RTX 5090PCIe Gen5 x16 | 128 GB | 64 | 256 GB | 4 TB | $1,316/mo | Configure 4× RTX 5090 |
GeForce cards come in servers of 1, 2 or 4 cards.
NVIDIA RTX 5090 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Blackwell
- Form factor
- PCIe Gen5 graphics card
- GPU memory
- 32 GB GDDR7, 512-bit
- Memory bandwidth
- 1,792 GB/s
- GPU-to-GPU link
- None (PCIe Gen5)
- Total graphics power
- 575 W
Tensor throughput (dense)
- FP4
- 1,676 TFLOPS
- FP8
- 419 TFLOPS
- FP16 / BF16
- 209.5 TFLOPS
- TF32
- 104.8 TFLOPS
- INT8
- 838 TOPS
- FP32
- 104.8 TFLOPS
Cores and media
- CUDA cores
- 21,760
- Tensor Cores
- 680, 5th generation
- RT Cores
- 170, 4th generation
- Video engines
- 3 NVENC, 2 NVDEC
Server, per GPU
- vCPU
- 16
- RAM
- 64 GB DDR5
- NVMe
- 1 TB
- GPUs per server
- 1, 2, 4
From NVIDIA’s RTX Blackwell architecture whitepaper, dense. NVIDIA’s headline “3,352 AI TOPS” is the sparse FP4 figure.
Which models fit on RTX 5090 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 4× | 2× | 1× |
| Llama 3.3 70B | — | 4× | 2× |
| gpt-oss-120b | — | — | 4× |
| DeepSeek-V4-Flash | — | — | — |
| Qwen3.5-397B-A17B | — | — | — |
| DeepSeek-R1 (671B) | — | — | — |
| Kimi K2.6 (1T) | — | — | — |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX 5090 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
RTX 5090 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
RTX 5090 questions.
More in the full FAQ.
RTX 5090 or RTX 4090?
The 5090 has 32 GB instead of 24 GB, 1.8× the memory bandwidth and FP4 support, for $60 more per month.
Can I run a 70B model on RTX 5090s?
In 4-bit a 70B model needs about 46 GB by our rule of thumb: two RTX 5090 hold it, split with tensor parallelism.
Is it good for Blender?
Yes: RT cores with OptiX make it one of the fastest cards for Cycles.
Build your RTX 5090 server.
Pick 1 to 4 GPUs and see the exact monthly price. $329 per GPU, the same at every size.
