Solutions · Image & video
GPU servers for image and video generation.
Run diffusion and video models, ComfyUI workflows and batch renders on GeForce and RTX PRO cards, by the month.
- 1× RTX 5090best fit
- 32 GBGPU memory
- $329per month
- From $329cheapest pick
What image and video generation needs from a GPU.
- 1
Fast GDDR7 memory
Diffusion and video models move large activations at every step: 1.8 TB/s on the RTX 5090 and 1.6 TB/s on the RTX PRO 6000.
- 2
Enough VRAM for video
Video models and high resolutions quickly exceed 24 GB. The RTX 5090 has 32 GB, the RTX PRO 6000 has 96 GB and the L40S 48 GB.
- 3
Low-precision tensor cores
Blackwell cards run FP8 and FP4, which recent diffusion pipelines use to speed up inference.
Our picks for image and video generation.
The first is the best fit; the other two trade speed, memory or ecosystem for price.
-
Best fit
1× NVIDIA RTX 5090
32 GB GDDR7 · 1.8 TB/s per GPU
- Interconnect
- Single GPU
- vCPU · RAM
- 16 · 64 GB
- NVMe
- 1 TB
$329/mo
$481 -
Alternative
1× NVIDIA RTX PRO 6000
96 GB GDDR7 ECC · 1.6 TB/s per GPU
- Interconnect
- Single GPU
- vCPU · RAM
- 24 · 180 GB
- NVMe
- 1.92 TB
$959/mo
$1,380 -
Alternative
1× NVIDIA L40S
48 GB GDDR6 ECC · 864 GB/s per GPU
- Interconnect
- Single GPU
- vCPU · RAM
- 12 · 96 GB
- NVMe
- 1 TB
$789/mo
$1,139
How big a server for your model?
Model weights in FP16/BF16 at ≈2.4 bytes per parameter with a margin. Resolution, video length and batch size add activation memory on top.
| Model size | FP16 / BF16 inference |
|---|---|
| 3Babout 7 GB | 1× RTX 5080$199/mo |
| 8Babout 19 GB | 1× RTX 4090$269/mo |
| 12Babout 29 GB | 1× RTX 5090$329/mo |
| 14Babout 34 GB | 2× RTX 4090$538/mo |
| 20Babout 48 GB | 2× RTX 5090$658/mo |
| 32Babout 77 GB | 1× RTX PRO 6000$959/mo |
Cheapest configuration that holds the model by our rule of thumb, 92% of GPU memory usable. Try other sizes
Tools people use, and what we recommend.
- ComfyUINode-based workflows for image and video diffusion models.
- Hugging Face DiffusersPython pipelines for diffusion models.
- PyTorch + xFormers / SDPAMemory-efficient attention for large resolutions.
- FFmpeg with NVENCHardware video encoding on NVIDIA cards.
Tips
- Run one worker per GPU and feed them from a queue: throughput scales with the number of cards.
- Keep model checkpoints on local NVMe; loading from the network slows every restart.
- Expose web UIs only through an SSH tunnel or an authenticated proxy.
Good to know
For generation, VRAM decides what you can run and memory bandwidth decides how fast. Several cards in one server run several jobs in parallel; one image generation job rarely spans several GPUs.
Related guides
- Your first hour on a new GPU serverYour server comes with its address and root SSH access; the software stack is yours to choose. This first hour covers login, hardware checks, updates and a sudo user.
- SSH access: keys, firewall, tunnels and file transferLock the server down to key-based SSH, keep long jobs running when you disconnect, and reach notebooks and dashboards through SSH instead of open ports.
- Docker with GPUs: NVIDIA and ROCm containersContainers keep CUDA, ROCm and framework versions off the host. The host needs a GPU driver, Docker and, for NVIDIA GPUs, the NVIDIA Container Toolkit.
Image & video questions.
Other workloads: LLM training, Fine-tuning, Inference & serving, 3D rendering & VFX, Research & HPC.
RTX 5090 or RTX PRO 6000?
The RTX 5090 is the best price per image when models fit in 32 GB. The RTX PRO 6000 has the same generation of GPU with 96 GB of ECC memory for large video models or several models loaded at once.
Can I run ComfyUI and reach it from my browser?
Yes: start it on the server and open it through an SSH tunnel, as described in the SSH guide. Avoid exposing it directly to the internet.
How many cards can a server have?
GeForce servers come with 1, 2 or 4 cards; RTX PRO 6000 and L40S servers with 1, 2, 4 or 8.
Start with 1× RTX 5090.
NVIDIA RTX 5090, 32 GB of GPU memory, for $329 a month.
