Pro & inference · Ada Lovelace
NVIDIA L40S servers
Ada Lovelace data-center GPU with 48 GB: cost-efficient inference and graphics.
- 48 GBGDDR6 ECC per GPU
- 864 GB/smemory bandwidth
- PCIe Gen4 x16GPU interconnect
- 1–8 GPUsper server
Price per GPU · monthly
$789/mo
−31%Market median $1,139≈ $1.08 per GPU-hour
Server size
Configure a L40S serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the L40S.
The L40S is NVIDIA’s Ada Lovelace data-center GPU: 48 GB of ECC GDDR6, FP8 tensor cores, RT cores and video engines, in a PCIe card built for servers. It is a cost-efficient choice for inference of models up to about 30B parameters, diffusion and video pipelines.
- 48 GB ECC
Twice the memory of the RTX 4090, with data-center reliability.
- FP8 tensor cores
Efficient low-precision inference.
- Graphics and video
RT cores plus hardware video encode and decode.
From 1 to 8 L40S GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× L40S | 48 GB | 12 | 96 GB | 1 TB | $789/mo | Configure 1× L40S | |
| 2× L40SPCIe Gen4 x16 | 96 GB | 24 | 192 GB | 2 TB | $1,578/mo | Configure 2× L40S | |
| 4× L40SPCIe Gen4 x16 | 192 GB | 48 | 384 GB | 4 TB | $3,156/mo | Configure 4× L40S | |
| 8× L40SPCIe Gen4 x16 | 384 GB | 96 | 768 GB | 8 TB | $6,312/mo | Configure 8× L40S |
NVIDIA L40S specifications.
Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- Ada Lovelace
- Form factor
- Dual-slot PCIe card, passive
- GPU memory
- 48 GB GDDR6 with ECC
- Memory bandwidth
- 864 GB/s
- GPU-to-GPU link
- None (PCIe Gen4 x16)
- Max power
- 350 W
- Multi-Instance GPU
- Not supported
Tensor throughput (dense)
- FP8
- 733 TFLOPS
- FP16 / BF16
- 362 TFLOPS
- TF32
- 183 TFLOPS
- INT8
- 733 TOPS
- FP32
- 91.6 TFLOPS
Cores and media
- CUDA cores
- 18,176
- Tensor Cores
- 568, 4th generation
- RT Cores
- 142, 3rd generation
- Video engines
- 3 NVENC, 3 NVDEC
Server, per GPU
- vCPU
- 12
- RAM
- 96 GB DDR5
- NVMe
- 1 TB
- GPUs per server
- 1, 2, 4, 8
NVIDIA lists both dense and sparse tensor figures; dense shown. Ada Lovelace has no FP4 tensor cores.
Which models fit on L40S servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 2× | 1× | 1× |
| Llama 3.3 70B | 4× | 2× | 2× |
| gpt-oss-120b | 8× | 4× | 2× |
| DeepSeek-V4-Flash | — | 8× | 8× |
| Qwen3.5-397B-A17B | — | — | 8× |
| DeepSeek-R1 (671B) | — | — | — |
| Kimi K2.6 (1T) | — | — | — |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest L40S server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
L40S prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
L40S questions.
More in the full FAQ.
L40S or RTX PRO 6000?
The RTX PRO 6000 is newer, with twice the memory and nearly twice the bandwidth; the L40S costs less. For models under ~30B parameters in FP8, the L40S is enough.
Does the L40S support FP8?
Yes, Ada Lovelace tensor cores support FP8.
How many per server?
1, 2, 4 or 8 cards.
Build your L40S server.
Pick 1 to 8 GPUs and see the exact monthly price. $789 per GPU, the same at every size.
