New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

16 open models · sized for our GPUs

Open models,
sized for your GPU.

How much GPU memory each model needs, and the cheapest server that runs it. Pick one at the Software step of the deploy page and it comes preinstalled.

  • 16language models
  • 5image and video models
  • vLLM · SGLangor Ollama
  • $0to preinstall

VRAM and cheapest server, model by model.

Official weights with vLLM or SGLang; the cheapest option may be FP8 or an Ollama build. Open a model for every option and GPU.

GPU memory needed per open model and cheapest server
ModelWeightsGPU memoryCheapest serverPer month
Qwen3.5 9BQwen · 9B 19 GB BF16 23 GB 1× RTX 5080FP8 $199/mo
gpt-oss-20bOpenAI · 21B MoE · 3.6B active 14 GB MXFP4 17 GB 1× RTX 4090MXFP4 $269/mo
Devstral Small 2Mistral AI · 24B 26 GB FP8 31 GB 1× RTX 4090Ollama $269/mo
Qwen3.8 27BQwen · 27B 56 GB BF16 67 GB 1× RTX 4090Ollama $269/mo
Gemma 4 31BGoogle · 31B 63 GB BF16 76 GB 1× RTX 5090Ollama $329/mo
Qwen3.6 35B-A3BQwen · 35B MoE · 3B active 72 GB BF16 86 GB 1× RTX 5090Ollama $329/mo
gpt-oss-120bOpenAI · 117B MoE · 5.1B active 65 GB MXFP4 78 GB 1× RTX PRO 6000MXFP4 $959/mo
Qwen3-Coder-NextQwen · 80B MoE · 3B active 159 GB BF16 191 GB 1× RTX PRO 6000Ollama $959/mo
Mistral Small 4Mistral AI · 119B MoE 121 GB FP8 145 GB 2× RTX PRO 6000FP8 $1,918/mo
DeepSeek-V4-FlashDeepSeek · 284B MoE · 13B active 167 GB FP4/FP8 200 GB 4× RTX PRO 6000FP4/FP8 $3,836/mo
GLM-5.3-FlashZ.ai · 320B MoE · 18B active 328 GB FP8 394 GB 8× RTX PRO 6000FP8 $7,672/mo
Qwen3.5 397B-A17BQwen · 397B MoE · 17B active 807 GB BF16 968 GB 8× RTX PRO 6000FP8 $7,672/mo
DeepSeek-R1-0528DeepSeek · 671B MoE · 37B active 689 GB FP8 827 GB 2× MI355XOllama $2,818/mo
Kimi K2.6Moonshot AI · 1T MoE · 32B active 595 GB INT4 714 GB 4× B300INT4 $16,076/mo
DeepSeek-V4-ProDeepSeek · 1.6T MoE · 49B active 893 GB FP4/FP8 1,072 GB 8× B200FP4/FP8 $27,512/mo
Kimi K3Moonshot AI · 2.8T MoE · 104B active 1561 GB MXFP4 1,873 GB 8× B300MXFP4 $32,152/mo

GPU memory = official weights × 1.2, with vLLM or SGLang. How much VRAM does an LLM need?

Image and video models for ComfyUI.

Installed with ComfyUI when you pick them at the Software step. Sizes are the full downloads.

Image and video models available with ComfyUI
ModelKindSizeLicence
Qwen-ImageQwen Image 20B · 58 GB Apache 2.0
Qwen-Image-Edit-2511Qwen Image editing 20B · 58 GB Apache 2.0
Wan 2.2 TI2V-5BWan-AI Video 5B · 34 GB Apache 2.0
Wan 2.2 T2V-A14BWan-AI Video 2 × 14B · 126 GB Apache 2.0
SDXL Base 1.0Stability AI Image 7 GB OpenRAIL++-M

For diffusion and video work, see GPU servers for image and video generation.

Open model FAQ.

Serving guide: vLLM on dedicated GPUs.

How do you estimate the GPU memory a model needs?

We take the size of the official weights and multiply it by 1.2 for the KV cache and activations, then keep 8% of GPU memory free. Long contexts and many parallel requests need more. It is the rule our deploy page uses to propose a server.

Can you preinstall a model on my server?

Yes, at no extra cost: pick the model at the Software step of the deploy page, with vLLM, SGLang or Ollama and Open WebUI. We download the official weights, and you are root on the server.

Why is there no Llama or FLUX model?

Their downloads require you to accept a licence on Hugging Face with your own account. We only preinstall models that anyone can download; you can install any other model yourself as root.

How current is this list?

Model sizes were read from the official repositories on Hugging Face on 23 September 2026. We update the list together with our price checks.

Pick a model. We size the server.

The deploy page proposes the cheapest server that holds your model and installs it for you. Paid in crypto, no KYC.

Customer area

Sign in to CryptGPU

Your orders, invoices and servers in one place.

New to CryptGPU?

Accounts are created at checkout: choose your server and software, then create your account in the Account step of your first order.

Deploy a server