New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Configure your GPU server.

Pick a GPU and a size, or start from the model you want to run. The monthly price, the market median and the full server spec update as you go.

GPU model

AI flagship

Pro & inference

Creator & consumer

GPUs per server

How we estimate GPU memory.

The “By model size” mode multiplies the parameter count by a number of bytes per parameter, then keeps 8% of GPU memory free.

Bytes of GPU memory per model parameter, by workload
WorkloadBytes / parameter70B modelCheapest fit
FP16 / BF16 inference ≈ 2.4 168 GB 1× MI355X$1,409/mo
FP8 inference ≈ 1.2 84 GB 1× RTX PRO 6000$959/mo
4-bit inference ≈ 0.65 46 GB 2× RTX 5090$658/mo
QLoRA fine-tune ≈ 0.75 53 GB 2× RTX 5090$658/mo
LoRA fine-tune (BF16) ≈ 3 210 GB 1× MI355X$1,409/mo
Full training (mixed precision, Adam) ≈ 18 1,260 GB 8× MI355X$11,272/mo

What the numbers include

  • Inference: weights (2, 1 or about 0.5 bytes per parameter) plus 20–30% for the KV cache.
  • LoRA and QLoRA: the frozen base model, adapters, gradients and activations.
  • Full training: weights, gradients and Adam optimizer states in mixed precision, plus a margin for activations.

Mixture-of-experts models count by their total parameters: every expert stays in memory. Long contexts and many concurrent requests need more memory for the KV cache.

Every configuration includes.

  • Dedicated GPUsNo time-slicing and no other tenant on your cards.
  • Root SSH accessInstall your drivers, CUDA or ROCm, containers and frameworks.
  • DDR5 RAM and local NVMevCPU, RAM and NVMe scale with the number of GPUs.
  • Same price per GPUAt every server size, from 1 to 8 GPUs.
  • Month to monthPrepaid monthly, renewed when you choose. No long-term contract.
  • Pay in cryptoBTC, ETH, USDT, USDC or LTC.

Configured? Here is what happens next.

  1. 1

    Keep your configuration

    Online checkout opens soon. The configuration link keeps the exact GPU, size and price.

  2. 2

    Pay in crypto

    One payment in BTC, ETH, USDT, USDC or LTC covers one month. The server starts once the payment is confirmed on-chain. Payment details

  3. 3

    Log in as root

    You receive the server address and root SSH access, then install your drivers, CUDA or ROCm and containers. Getting started

Configuration questions.

Anything else? Ask us.

Is the price per GPU the same at every size?

Yes. A server with n GPUs costs n times the price of one GPU, from 1 to 8 GPUs. vCPU, RAM and NVMe scale in the same proportion.

Can I mix GPU models in one server?

No. Each server has one GPU model. To run two models, for example H200 for serving and RTX 5090 for image generation, configure two servers.

Can I add GPUs to my server later?

Order the new size as a new server and move your data; the old server keeps its term until it ends. Moving data between servers is your responsibility.

Why can’t I pick 8 GeForce cards?

GeForce servers (RTX 5090, 4090, 5080) come with 1, 2 or 4 cards. Data-center GPUs and the RTX PRO 6000 and L40S come with 1, 2, 4 or 8.

How accurate is the model sizing?

It is a rule of thumb: bytes per parameter for the chosen precision or training method, with headroom for the KV cache and activations, and 92% of GPU memory counted as usable. Very long contexts, large batches and some frameworks need more; test with your own workload.

How do I order this configuration?

Online checkout opens soon. Copy the configuration link to keep your exact server and price, or contact us with it if you have questions.