Creator & consumer · Ada Lovelace
NVIDIA RTX 4090 servers
Ada Lovelace GeForce with 24 GB: the budget workhorse for AI and rendering.
- 24 GBGDDR6X per GPU
- 1 TB/smemory bandwidth
- PCIe Gen4 x16GPU interconnect
- 1–4 GPUsper server
Price per GPU · monthly
$269/mo
−30%Market median $387≈ $0.37 per GPU-hour
Below every on-demand list price we found
Server size
Configure a RTX 4090 serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the RTX 4090.
The GeForce RTX 4090 remains one of the most used GPUs for affordable AI work: 24 GB of GDDR6X, FP8-capable tensor cores and RT cores. It suits Stable Diffusion-class models, 7B–14B LLMs in 16-bit, larger ones quantized, and GPU rendering.
- 24 GB GDDR6X
Enough for 13B models in FP8 or 30B in 4-bit.
- Low monthly price
$269 per card per month.
- Mature software support
Widely used, well-supported by AI and rendering tools.
From 1 to 4 RTX 4090 GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× RTX 4090 | 24 GB | 12 | 64 GB | 1 TB | $269/mo | Configure 1× RTX 4090 | |
| 2× RTX 4090PCIe Gen4 x16 | 48 GB | 24 | 128 GB | 2 TB | $538/mo | Configure 2× RTX 4090 | |
| 4× RTX 4090PCIe Gen4 x16 | 96 GB | 48 | 256 GB | 4 TB | $1,076/mo | Configure 4× RTX 4090 |
GeForce cards come in servers of 1, 2 or 4 cards.
NVIDIA RTX 4090 specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Ada Lovelace, TSMC 4N
- Form factor
- PCIe Gen4 graphics card
- GPU memory
- 24 GB GDDR6X, 384-bit
- Memory bandwidth
- 1,008 GB/s
- GPU-to-GPU link
- None (PCIe Gen4)
- Graphics card power
- 450 W
Tensor throughput (dense)
- FP8
- 330 TFLOPS
- FP16 / BF16
- 165 TFLOPS
- TF32
- 82.6 TFLOPS
- INT8
- 661 TOPS
- FP32
- 82.6 TFLOPS
Cores and media
- CUDA cores
- 16,384
- Tensor Cores
- 512, 4th generation
- RT Cores
- 128, 3rd generation
- Video engines
- 2 NVENC, 1 NVDEC
Server, per GPU
- vCPU
- 12
- RAM
- 64 GB DDR5
- NVMe
- 1 TB
- GPUs per server
- 1, 2, 4
From NVIDIA’s Ada architecture whitepaper, dense. Ada Lovelace has no FP4 tensor cores.
Which models fit on RTX 4090 servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 4× | 2× | 1× |
| Llama 3.3 70B | — | 4× | 4× |
| gpt-oss-120b | — | — | 4× |
| DeepSeek-V4-Flash | — | — | — |
| Qwen3.5-397B-A17B | — | — | — |
| DeepSeek-R1 (671B) | — | — | — |
| Kimi K2.6 (1T) | — | — | — |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest RTX 4090 server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
RTX 4090 prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
RTX 4090 questions.
More in the full FAQ.
What can I run on one RTX 4090?
Image models like SDXL-class and FLUX-class in FP16, and LLMs up to about 14B in FP8 or 30B in 4-bit by our rule of thumb.
Why choose it over the RTX 5090?
It costs $60 less per month. The 5090 is faster and has 8 GB more memory.
How many per server?
1, 2 or 4 cards.
Build your RTX 4090 server.
Pick 1 to 4 GPUs and see the exact monthly price. $269 per GPU, the same at every size.
