GPU comparison · Ada Lovelace vs Hopper
NVIDIA L40S vs H100.
A PCIe inference card against an SXM training GPU. The L40S (Ada Lovelace) has 48 GB of GDDR6 with ECC, FP8, RT cores and video engines at less than half the price of the H100, which has 80 GB of HBM3 at 3.35 TB/s and NVLink.
- 48 vs 80 GBGPU memory
- 0.86 vs 3.35 TB/smemory bandwidth
- $789 vs $1,779per GPU, per month
L40S vs H100: which one to rent.
Choose the L40S if
- You serve models up to about 30B parameters in FP8, or run diffusion and video pipelines.
- You need graphics as well as AI: RT cores and hardware video encode and decode.
- Price matters: $789 per GPU against $1,779 a month.
Choose the H100 if
- You train or fine-tune across GPUs: NVLink 4 and 3.35 TB/s of HBM3.
- You serve larger models: 80 GB per GPU and nearly 4× the bandwidth.
- You need FP64 for simulation.
L40S vs H100 specs, side by side.
Figures from NVIDIA’s published specifications, per GPU. Tensor figures are peak dense throughput.
| Per GPU | NVIDIA L40S | NVIDIA H100 | H100 / L40S |
|---|---|---|---|
| Architecture | Ada Lovelace | Hopper, TSMC 4N | — |
| Form factor | Dual-slot PCIe card, passive | SXM5 module | — |
| GPU memory | 48 GB GDDR6 with ECC | 80 GB HBM3, 5,120-bit | 1.67× |
| Memory bandwidth | 864 GB/s | 3.35 TB/s | 3.9× |
| GPU-to-GPU link | None (PCIe Gen4 x16) | NVLink 4, 900 GB/s per GPU | — |
| FP4 tensor (dense) | Not supported | Not supported | — |
| FP8 tensor (dense) | 733 TFLOPS | 1,979 TFLOPS | 2.7× |
| FP16 / BF16 tensor (dense) | 362 TFLOPS | 989 TFLOPS | 2.7× |
| FP64 | — | 34 / 67 TFLOPS | — |
| Max power | 350 W | Up to 700 W, configurable | — |
| GPUs per server | 1, 2, 4, 8 | 1, 2, 4, 8 | — |
| Per GPU in our servers | 12 vCPU · 96 GB RAM · 1 TB | 20 vCPU · 200 GB RAM · 2 TB | — |
Sources: NVIDIA L40S · NVIDIA H100. Full sheets: L40S · H100
L40S vs H100 rental price.
Our monthly prices, set 30% below the market median of public on-demand prices and paid in crypto. Hourly figures are the monthly price divided by 730 hours.
| CryptGPU | NVIDIA L40S | NVIDIA H100 | H100 / L40S |
|---|---|---|---|
| Price per GPU, per month | $789 | $1,779 | 2.3× |
| Equivalent per GPU-hour | $1.08 | $2.44 | — |
| Per GB of GPU memory, per month | $16.44 | $22.24 | 1.35× |
| Per PFLOPS of dense FP8, per month | $1,076 | $899 | 0.84× |
| Largest server | 8× · $6,312/mo | 8× · $14,232/mo | — |
| Market median, per GPU | $1,139 | $2,548 | — |
| Below the median | −31% | −30% | — |
Same price per GPU at every server size, no hourly metering. How we compare prices · Configure an L40S server · Configure a H100 server
Which models fit on each.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache, 92% of GPU memory usable.
| Model (total parameters) | L40S · FP8 | H100 · FP8 | L40S · 4-bit | H100 · 4-bit |
|---|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× | 1× |
| Llama 3.3 70B | 2× | 2× | 2× | 1× |
| gpt-oss-120b | 4× | 2× | 2× | 2× |
| DeepSeek-V4-Flash | 8× | 8× | 8× | 4× |
| Qwen3.5-397B-A17B | — | 8× | 8× | 4× |
| DeepSeek-R1 (671B) | — | — | — | 8× |
| Kimi K2.6 (1T) | — | — | — | — |
FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter. — = larger than the biggest server of that GPU. Size another model · How much VRAM does an LLM need?
L40S vs H100 FAQ.
More on each GPU: NVIDIA L40S · NVIDIA H100.
Is the L40S good for LLM inference?
For models that fit in 48 GB, yes: FP8 tensor cores and ECC memory at a low price. With 864 GB/s of bandwidth, it generates tokens more slowly than HBM GPUs on large models.
How much cheaper is the L40S?
$990 less per GPU per month: $789 against $1,779.
Can the L40S train models?
It can fine-tune smaller models, but it has no NVLink and less bandwidth: for multi-GPU training, the H100 or newer HBM GPUs are the right choice.
Other GPU comparisons.
Rent the L40S or the H100.
Dedicated servers with 1 to 8 GPUs, one monthly price, paid in crypto. Online in under 10 minutes, no KYC.
