New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Creator & consumer · Blackwell

NVIDIA RTX 5080 servers

Blackwell GeForce with 16 GB: the entry GPU server.

  • 16 GBGDDR7 per GPU
  • 960 GB/smemory bandwidth
  • PCIe Gen5 x16GPU interconnect
  • 1–4 GPUsper server

Price per GPU · monthly

$199/mo

−33%

Market median $299≈ $0.27 per GPU-hour

Below every on-demand list price we found

Server size

Configure a RTX 5080 server

USD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC

Why the RTX 5080.

The GeForce RTX 5080 is the most affordable GPU in our range: Blackwell architecture, 16 GB of GDDR7 at 960 GB/s and FP4/FP8 tensor cores. It suits small models, CI pipelines for GPU code, development environments and video encoding.

  • From $199 per month

    The lowest entry price in our range.

  • Blackwell tensor cores

    FP4 and FP8 support on a budget.

  • GDDR7 at 960 GB/s

    Fast memory for its class.

From 1 to 4 RTX 5080 GPUs.

Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.

NVIDIA RTX 5080 server sizes and monthly prices
Server GPU memory vCPU RAM NVMe Market median CryptGPU Action
1× RTX 5080 16 GB 8 32 GB 0.5 TB $299 $199/mo Configure 1× RTX 5080
2× RTX 5080PCIe Gen5 x16 32 GB 16 64 GB 1 TB $598 $398/mo Configure 2× RTX 5080
4× RTX 5080PCIe Gen5 x16 64 GB 32 128 GB 2 TB $1,196 $796/mo Configure 4× RTX 5080

GeForce cards come in servers of 1, 2 or 4 cards.

NVIDIA RTX 5080 specifications.

Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.

GPU and memory

Architecture
Blackwell
Form factor
PCIe Gen5 graphics card
GPU memory
16 GB GDDR7, 256-bit
Memory bandwidth
960 GB/s
GPU-to-GPU link
None (PCIe Gen5)
Total graphics power
360 W

Tensor throughput (dense)

FP4
900 TFLOPS
FP8
225 TFLOPS
FP16 / BF16
112.6 TFLOPS
TF32
56.3 TFLOPS
INT8
450 TOPS
FP32
56.3 TFLOPS

Cores and media

CUDA cores
10,752
Tensor Cores
336, 5th generation
RT Cores
84, 4th generation
Video engines
2 NVENC, 2 NVDEC

Server, per GPU

vCPU
8
RAM
32 GB DDR5
NVMe
0.5 TB
GPUs per server
1, 2, 4

From NVIDIA’s RTX Blackwell architecture whitepaper, dense. NVIDIA’s headline “1,801 AI TOPS” is the sparse FP4 figure.

Which models fit on RTX 5080 servers.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.

Number of RTX 5080 GPUs needed per model and precision
Model (total parameters)FP16 / BF16FP84-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX 5080 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.

RTX 5080 prices across the market.

Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.

  • CryptGPUmonthly, dedicated$199$0.27/h
  • Lowest list priceon-demand, 2 providers$285$0.39/h
  • Market medianindicative reference$299$0.41/h
  • Highest list priceon-demand$313$0.43/h

Per GPU, 730 hours a month. No provider publishes an on-demand price for this GPU yet: its market reference is the median of the public prices we found. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare

RTX 5080 questions.

More in the full FAQ.

What fits in 16 GB?

By our rule of thumb, models up to about 12B parameters in FP8 or 20B in 4-bit, and most image models at standard resolutions.

Is it good for development?

Yes: a dedicated CUDA GPU with root access for building, testing and CI at the lowest monthly price.

How many per server?

1, 2 or 4 cards.

Build your RTX 5080 server.

Pick 1 to 4 GPUs and see the exact monthly price. $199 per GPU, the same at every size.