New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Solutions · Image & video

GPU servers for image and video generation.

Run diffusion and video models, ComfyUI workflows and batch renders on GeForce and RTX PRO cards, by the month.

  • 1× RTX 5090best fit
  • 32 GBGPU memory
  • $329per month
  • From $329cheapest pick

What image and video generation needs from a GPU.

  1. 1

    Fast GDDR7 memory

    Diffusion and video models move large activations at every step: 1.8 TB/s on the RTX 5090 and 1.6 TB/s on the RTX PRO 6000.

  2. 2

    Enough VRAM for video

    Video models and high resolutions quickly exceed 24 GB. The RTX 5090 has 32 GB, the RTX PRO 6000 has 96 GB and the L40S 48 GB.

  3. 3

    Low-precision tensor cores

    Blackwell cards run FP8 and FP4, which recent diffusion pipelines use to speed up inference.

Our picks for image and video generation.

The first is the best fit; the other two trade speed, memory or ecosystem for price.

  • Best fit

    1× NVIDIA RTX 5090

    32 GB GDDR7 · 1.8 TB/s per GPU

    Interconnect
    Single GPU
    vCPU · RAM
    16 · 64 GB
    NVMe
    1 TB

    $329/mo$481

  • Alternative

    1× NVIDIA RTX PRO 6000

    96 GB GDDR7 ECC · 1.6 TB/s per GPU

    Interconnect
    Single GPU
    vCPU · RAM
    24 · 180 GB
    NVMe
    1.92 TB

    $959/mo$1,380

  • Alternative

    1× NVIDIA L40S

    48 GB GDDR6 ECC · 864 GB/s per GPU

    Interconnect
    Single GPU
    vCPU · RAM
    12 · 96 GB
    NVMe
    1 TB

    $789/mo$1,139

How big a server for your model?

Model weights in FP16/BF16 at ≈2.4 bytes per parameter with a margin. Resolution, video length and batch size add activation memory on top.

Cheapest server by model size for image and video generation
Model sizeFP16 / BF16 inference
3Babout 7 GB 1× RTX 5080$199/mo
8Babout 19 GB 1× RTX 4090$269/mo
12Babout 29 GB 1× RTX 5090$329/mo
14Babout 34 GB 2× RTX 4090$538/mo
20Babout 48 GB 2× RTX 5090$658/mo
32Babout 77 GB 1× RTX PRO 6000$959/mo

Cheapest configuration that holds the model by our rule of thumb, 92% of GPU memory usable. Try other sizes

Tools people use, and what we recommend.

  • ComfyUINode-based workflows for image and video diffusion models.
  • Hugging Face DiffusersPython pipelines for diffusion models.
  • PyTorch + xFormers / SDPAMemory-efficient attention for large resolutions.
  • FFmpeg with NVENCHardware video encoding on NVIDIA cards.

Tips

  • Run one worker per GPU and feed them from a queue: throughput scales with the number of cards.
  • Keep model checkpoints on local NVMe; loading from the network slows every restart.
  • Expose web UIs only through an SSH tunnel or an authenticated proxy.

Good to know

For generation, VRAM decides what you can run and memory bandwidth decides how fast. Several cards in one server run several jobs in parallel; one image generation job rarely spans several GPUs.

Image & video questions.

Other workloads: LLM training, Fine-tuning, Inference & serving, 3D rendering & VFX, Research & HPC.

RTX 5090 or RTX PRO 6000?

The RTX 5090 is the best price per image when models fit in 32 GB. The RTX PRO 6000 has the same generation of GPU with 96 GB of ECC memory for large video models or several models loaded at once.

Can I run ComfyUI and reach it from my browser?

Yes: start it on the server and open it through an SSH tunnel, as described in the SSH guide. Avoid exposing it directly to the internet.

How many cards can a server have?

GeForce servers come with 1, 2 or 4 cards; RTX PRO 6000 and L40S servers with 1, 2, 4 or 8.

Start with 1× RTX 5090.

NVIDIA RTX 5090, 32 GB of GPU memory, for $329 a month.