Creator & consumer · Blackwell
NVIDIA RTX 5080 servers
Blackwell GeForce with 16 GB: the entry GPU server.
- 16 GBGDDR7 per GPU
- 960 GB/smemory bandwidth
- PCIe Gen5 x16GPU interconnect
- 1–4 GPUsper server
Price per GPU · monthly
$199/mo
−33%Market median $299≈ $0.27 per GPU-hour
Below every on-demand list price we found
Server size
Configure a RTX 5080 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the RTX 5080.
The GeForce RTX 5080 is the most affordable GPU in our range: Blackwell architecture, 16 GB of GDDR7 at 960 GB/s and FP4/FP8 tensor cores. It suits small models, CI pipelines for GPU code, development environments and video encoding.
- From $199 per month
The lowest entry price in our range.
- Blackwell tensor cores
FP4 and FP8 support on a budget.
- GDDR7 at 960 GB/s
Fast memory for its class.
From 1 to 4 RTX 5080 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× RTX 5080 | 16 GB | 8 | 32 GB | 0.5 TB | $199/mo | Configure 1× RTX 5080 | |
| 2× RTX 5080PCIe Gen5 x16 | 32 GB | 16 | 64 GB | 1 TB | $398/mo | Configure 2× RTX 5080 | |
| 4× RTX 5080PCIe Gen5 x16 | 64 GB | 32 | 128 GB | 2 TB | $796/mo | Configure 4× RTX 5080 |
GeForce cards come in servers of 1, 2 or 4 cards.
NVIDIA RTX 5080 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Blackwell
- Form factor
- PCIe Gen5 graphics card
- GPU memory
- 16 GB GDDR7, 256-bit
- Memory bandwidth
- 960 GB/s
- GPU-to-GPU link
- None (PCIe Gen5)
- Total graphics power
- 360 W
Tensor throughput (dense)
- FP4
- 900 TFLOPS
- FP8
- 225 TFLOPS
- FP16 / BF16
- 112.6 TFLOPS
- TF32
- 56.3 TFLOPS
- INT8
- 450 TOPS
- FP32
- 56.3 TFLOPS
Cores and media
- CUDA cores
- 10,752
- Tensor Cores
- 336, 5th generation
- RT Cores
- 84, 4th generation
- Video engines
- 2 NVENC, 2 NVDEC
Server, per GPU
- vCPU
- 8
- RAM
- 32 GB DDR5
- NVMe
- 0.5 TB
- GPUs per server
- 1, 2, 4
From NVIDIA’s RTX Blackwell architecture whitepaper, dense. NVIDIA’s headline “1,801 AI TOPS” is the sparse FP4 figure.
Which models fit on RTX 5080 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 2× | 1× | 1× |
| Gemma 4 31B | — | 4× | 2× |
| Llama 3.3 70B | — | — | 4× |
| gpt-oss-120b | — | — | — |
| DeepSeek-V4-Flash | — | — | — |
| Qwen3.5-397B-A17B | — | — | — |
| DeepSeek-R1 (671B) | — | — | — |
| Kimi K2.6 (1T) | — | — | — |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX 5080 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
RTX 5080 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. No provider publishes an on-demand price for this GPU yet: its market reference is the median of the public prices we found. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
RTX 5080 questions.
More in the full FAQ.
What fits in 16 GB?
By our rule of thumb, models up to about 12B parameters in FP8 or 20B in 4-bit, and most image models at standard resolutions.
Is it good for development?
Yes: a dedicated CUDA GPU with root access for building, testing and CI at the lowest monthly price.
How many per server?
1, 2 or 4 cards.
Build your RTX 5080 server.
Pick 1 to 4 GPUs and see the exact monthly price. $199 per GPU, the same at every size.
