New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

GPU comparison · Ada Lovelace vs Hopper

NVIDIA L40S vs H100.

A PCIe inference card against an SXM training GPU. The L40S (Ada Lovelace) has 48 GB of GDDR6 with ECC, FP8, RT cores and video engines at less than half the price of the H100, which has 80 GB of HBM3 at 3.35 TB/s and NVLink.

  • 48 vs 80 GBGPU memory
  • 0.86 vs 3.35 TB/smemory bandwidth
  • $789 vs $1,779per GPU, per month

L40S vs H100: which one to rent.

Choose the L40S if

  • You serve models up to about 30B parameters in FP8, or run diffusion and video pipelines.
  • You need graphics as well as AI: RT cores and hardware video encode and decode.
  • Price matters: $789 per GPU against $1,779 a month.

NVIDIA L40S servers from $789/mo

Choose the H100 if

  • You train or fine-tune across GPUs: NVLink 4 and 3.35 TB/s of HBM3.
  • You serve larger models: 80 GB per GPU and nearly 4× the bandwidth.
  • You need FP64 for simulation.

NVIDIA H100 servers from $1,779/mo

L40S vs H100 specs, side by side.

Figures from NVIDIA’s published specifications, per GPU. Tensor figures are peak dense throughput.

L40S vs H100 specifications
Per GPUNVIDIA L40SNVIDIA H100H100 / L40S
ArchitectureAda LovelaceHopper, TSMC 4N
Form factorDual-slot PCIe card, passiveSXM5 module
GPU memory48 GB GDDR6 with ECC80 GB HBM3, 5,120-bit1.67×
Memory bandwidth864 GB/s3.35 TB/s3.9×
GPU-to-GPU linkNone (PCIe Gen4 x16)NVLink 4, 900 GB/s per GPU
FP4 tensor (dense)Not supportedNot supported
FP8 tensor (dense)733 TFLOPS1,979 TFLOPS2.7×
FP16 / BF16 tensor (dense)362 TFLOPS989 TFLOPS2.7×
FP6434 / 67 TFLOPS
Max power350 WUp to 700 W, configurable
GPUs per server1, 2, 4, 81, 2, 4, 8
Per GPU in our servers12 vCPU · 96 GB RAM · 1 TB20 vCPU · 200 GB RAM · 2 TB

Sources: NVIDIA L40S · NVIDIA H100. Full sheets: L40S · H100

L40S vs H100 rental price.

Our monthly prices, set 30% below the market median of public on-demand prices and paid in crypto. Hourly figures are the monthly price divided by 730 hours.

L40S vs H100 rental prices
CryptGPUNVIDIA L40SNVIDIA H100H100 / L40S
Price per GPU, per month$789$1,7792.3×
Equivalent per GPU-hour$1.08$2.44
Per GB of GPU memory, per month$16.44$22.241.35×
Per PFLOPS of dense FP8, per month$1,076$8990.84×
Largest server8× · $6,312/mo8× · $14,232/mo
Market median, per GPU$1,139$2,548
Below the median−31%−30%

Same price per GPU at every server size, no hourly metering. How we compare prices · Configure an L40S server · Configure a H100 server

Which models fit on each.

Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache, 92% of GPU memory usable.

Number of GPUs needed per model on L40S vs H100
Model (total parameters)L40S · FP8H100 · FP8L40S · 4-bitH100 · 4-bit
Qwen3.5-9B
Gemma 4 31B
Llama 3.3 70B
gpt-oss-120b
DeepSeek-V4-Flash
Qwen3.5-397B-A17B
DeepSeek-R1 (671B)
Kimi K2.6 (1T)

FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter. — = larger than the biggest server of that GPU. Size another model · How much VRAM does an LLM need?

L40S vs H100 FAQ.

More on each GPU: NVIDIA L40S · NVIDIA H100.

Is the L40S good for LLM inference?

For models that fit in 48 GB, yes: FP8 tensor cores and ECC memory at a low price. With 864 GB/s of bandwidth, it generates tokens more slowly than HBM GPUs on large models.

How much cheaper is the L40S?

$990 less per GPU per month: $789 against $1,779.

Can the L40S train models?

It can fine-tune smaller models, but it has no NVLink and less bandwidth: for multi-GPU training, the H100 or newer HBM GPUs are the right choice.

Rent the L40S or the H100.

Dedicated servers with 1 to 8 GPUs, one monthly price, paid in crypto. Online in under 10 minutes, no KYC.

Customer area

Sign in to CryptGPU

Your orders, invoices and servers in one place.

New to CryptGPU?

Accounts are created at checkout: choose your server and software, then create your account in the Account step of your first order.

Deploy a server