New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Pro & inference · Ada Lovelace

NVIDIA L40S servers

Ada Lovelace data-center GPU with 48 GB: cost-efficient inference and graphics.

  • 48 GBGDDR6 ECC per GPU
  • 864 GB/smemory bandwidth
  • PCIe Gen4 x16GPU interconnect
  • 1–8 GPUsper server

Price per GPU · monthly

$789/mo

−31%

Market median $1,139≈ $1.08 per GPU-hour

Server size

Configure a L40S server

USD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC

Why the L40S.

The L40S is NVIDIA’s Ada Lovelace data-center GPU: 48 GB of ECC GDDR6, FP8 tensor cores, RT cores and video engines, in a PCIe card built for servers. It is a cost-efficient choice for inference of models up to about 30B parameters, diffusion and video pipelines.

  • 48 GB ECC

    Twice the memory of the RTX 4090, with data-center reliability.

  • FP8 tensor cores

    Efficient low-precision inference.

  • Graphics and video

    RT cores plus hardware video encode and decode.

From 1 to 8 L40S GPUs.

Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.

NVIDIA L40S server sizes and monthly prices
Server GPU memory vCPU RAM NVMe Market median CryptGPU Action
1× L40S 48 GB 12 96 GB 1 TB $1,139 $789/mo Configure 1× L40S
2× L40SPCIe Gen4 x16 96 GB 24 192 GB 2 TB $2,278 $1,578/mo Configure 2× L40S
4× L40SPCIe Gen4 x16 192 GB 48 384 GB 4 TB $4,556 $3,156/mo Configure 4× L40S
8× L40SPCIe Gen4 x16 384 GB 96 768 GB 8 TB $9,112 $6,312/mo Configure 8× L40S

NVIDIA L40S specifications.

Figures from NVIDIA’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.

GPU and memory

Architecture
Ada Lovelace
Form factor
Dual-slot PCIe card, passive
GPU memory
48 GB GDDR6 with ECC
Memory bandwidth
864 GB/s
GPU-to-GPU link
None (PCIe Gen4 x16)
Max power
350 W
Multi-Instance GPU
Not supported

Tensor throughput (dense)

FP8
733 TFLOPS
FP16 / BF16
362 TFLOPS
TF32
183 TFLOPS
INT8
733 TOPS
FP32
91.6 TFLOPS

Cores and media

CUDA cores
18,176
Tensor Cores
568, 4th generation
RT Cores
142, 3rd generation
Video engines
3 NVENC, 3 NVDEC

Server, per GPU

vCPU
12
RAM
96 GB DDR5
NVMe
1 TB
GPUs per server
1, 2, 4, 8

NVIDIA lists both dense and sparse tensor figures; dense shown. Ada Lovelace has no FP4 tensor cores.

Which models fit on L40S servers.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.

Number of L40S GPUs needed per model and precision
Model (total parameters)FP16 / BF16FP84-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest L40S server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.

L40S prices across the market.

Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.

  • CryptGPUmonthly, dedicated$789$1.08/h
  • Lowest list priceon-demand, 10 providers$472$0.65/h
  • Market medianon-demand$1,139$1.56/h
  • Highest list priceon-demand$1,643$2.25/h
  • AWSg6e.xlarge · reference$1,359$1.86/h
  • Oracle Cloudlist price · reference$2,555$3.50/h

Per GPU, 730 hours a month. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare

L40S questions.

More in the full FAQ.

L40S or RTX PRO 6000?

The RTX PRO 6000 is newer, with twice the memory and nearly twice the bandwidth; the L40S costs less. For models under ~30B parameters in FP8, the L40S is enough.

Does the L40S support FP8?

Yes, Ada Lovelace tensor cores support FP8.

How many per server?

1, 2, 4 or 8 cards.

Build your L40S server.

Pick 1 to 8 GPUs and see the exact monthly price. $789 per GPU, the same at every size.