New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Creator & consumer · Ada Lovelace

NVIDIA RTX 4090 servers

Ada Lovelace GeForce with 24 GB: the budget workhorse for AI and rendering.

  • 24 GBGDDR6X per GPU
  • 1 TB/smemory bandwidth
  • PCIe Gen4 x16GPU interconnect
  • 1–4 GPUsper server

Price per GPU · monthly

$269/mo

−30%

Market median $387≈ $0.37 per GPU-hour

Below every on-demand list price we found

Server size

Configure a RTX 4090 server

USD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC

Why the RTX 4090.

The GeForce RTX 4090 remains one of the most used GPUs for affordable AI work: 24 GB of GDDR6X, FP8-capable tensor cores and RT cores. It suits Stable Diffusion-class models, 7B–14B LLMs in 16-bit, larger ones quantized, and GPU rendering.

  • 24 GB GDDR6X

    Enough for 13B models in FP8 or 30B in 4-bit.

  • Low monthly price

    $269 per card per month.

  • Mature software support

    Widely used, well-supported by AI and rendering tools.

From 1 to 4 RTX 4090 GPUs.

Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.

NVIDIA RTX 4090 server sizes and monthly prices
Server GPU memory vCPU RAM NVMe Market median CryptGPU Action
1× RTX 4090 24 GB 12 64 GB 1 TB $387 $269/mo Configure 1× RTX 4090
2× RTX 4090PCIe Gen4 x16 48 GB 24 128 GB 2 TB $774 $538/mo Configure 2× RTX 4090
4× RTX 4090PCIe Gen4 x16 96 GB 48 256 GB 4 TB $1,548 $1,076/mo Configure 4× RTX 4090

GeForce cards come in servers of 1, 2 or 4 cards.

NVIDIA RTX 4090 specifications.

Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.

GPU and memory

Architecture
Ada Lovelace, TSMC 4N
Form factor
PCIe Gen4 graphics card
GPU memory
24 GB GDDR6X, 384-bit
Memory bandwidth
1,008 GB/s
GPU-to-GPU link
None (PCIe Gen4)
Graphics card power
450 W

Tensor throughput (dense)

FP8
330 TFLOPS
FP16 / BF16
165 TFLOPS
TF32
82.6 TFLOPS
INT8
661 TOPS
FP32
82.6 TFLOPS

Cores and media

CUDA cores
16,384
Tensor Cores
512, 4th generation
RT Cores
128, 3rd generation
Video engines
2 NVENC, 1 NVDEC

Server, per GPU

vCPU
12
RAM
64 GB DDR5
NVMe
1 TB
GPUs per server
1, 2, 4

From NVIDIA’s Ada architecture whitepaper, dense. Ada Lovelace has no FP4 tensor cores.

Which models fit on RTX 4090 servers.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.

Number of RTX 4090 GPUs needed per model and precision
Model (total parameters)FP16 / BF16FP84-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX 4090 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.

RTX 4090 prices across the market.

Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.

  • CryptGPUmonthly, dedicated$269$0.37/h
  • Lowest list priceon-demand, 16 providers$283$0.39/h
  • Market medianon-demand$387$0.53/h
  • Highest list priceon-demand$887$1.22/h

Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare

RTX 4090 questions.

More in the full FAQ.

What can I run on one RTX 4090?

Image models like SDXL-class and FLUX-class in FP16, and LLMs up to about 14B in FP8 or 30B in 4-bit by our rule of thumb.

Why choose it over the RTX 5090?

It costs $60 less per month. The 5090 is faster and has 8 GB more memory.

How many per server?

1, 2 or 4 cards.

Build your RTX 4090 server.

Pick 1 to 4 GPUs and see the exact monthly price. $269 per GPU, the same at every size.