Open model · DeepSeek · MIT
Run DeepSeek-R1-0528
on your own GPUs.
DeepSeek’s 671B reasoning model, May 2025 update. 671B MoE · 37B active parameters, 689 GB of official FP8 weights.
- 827 GBGPU memory, FP8
- 671B MoE · 37B activeparameters
- 2× MI355Xcheapest server, Ollama
- $2,818per month
How much VRAM DeepSeek-R1-0528 needs.
Weights × 1.2 for the KV cache and activations, with 92% of GPU memory usable: the rule our deploy page uses. Long contexts and many parallel requests need more.
| Serving option | GPU memory | Cheapest server | Per month | Action |
|---|---|---|---|---|
| vLLM or SGLang, official FP8 weights | 827 GB | 4× AMD MI355X | $5,636/mo | Deploy |
| Ollama + Open WebUI, deepseek-r1:671b (quantized) | 485 GB | 2× AMD MI355X | $2,818/mo | Deploy |
Weights: deepseek-ai/DeepSeek-R1-0528 · 689 GB, FP8 · Ollama: deepseek-r1:671b, 404 GB. How we estimate GPU memory
Which GPUs run DeepSeek-R1-0528.
Smallest server of each GPU that holds the model, and its monthly price.
| GPU | FP8 | Ollama | Price, FP8 |
|---|---|---|---|
| NVIDIA B300288 GB HBM3e | 4× | 2× | $16,076/mo |
| NVIDIA B200180 GB HBM3e | 8× | 4× | $27,512/mo |
| AMD MI355X288 GB HBM3E | 4× | 2× | $5,636/mo |
| NVIDIA H200141 GB HBM3e | 8× | 4× | $18,232/mo |
| NVIDIA H10080 GB HBM3 | — | 8× | — |
| NVIDIA RTX PRO 600096 GB GDDR7 ECC | — | 8× | — |
| NVIDIA L40S48 GB GDDR6 ECC | — | — | — |
| NVIDIA RTX 509032 GB GDDR7 | — | — | — |
| NVIDIA RTX 409024 GB GDDR6X | — | — | — |
| NVIDIA RTX 508016 GB GDDR7 | — | — | — |
n/a = ROCm support for this model is not confirmed with vLLM or SGLang. — = larger than the biggest server of that GPU. Size another model
Deploy DeepSeek-R1-0528 preinstalled.
Pick the model at the Software step of the deploy page: we install it with the engine you choose, at no extra cost.
- 1
Pick the server
The deploy page proposes the cheapest server that holds the model, and switches when you change the precision or the engine.
- 2
Pick the engine
vLLM or SGLang serve an OpenAI-compatible API; Ollama with Open WebUI gives a private chat in the browser.
- 3
Pay and log in
Pay the month in BTC, ETH, USDT, XMR or LTC, no KYC. Your server is online in under 10 minutes after confirmation, with root SSH access.
Guides: serving LLMs with vLLM · GPU servers for LLM inference · first steps on your server
DeepSeek-R1-0528 FAQ.
Every model: open models and their VRAM.
How much VRAM does DeepSeek-R1-0528 need?
About 827 GB of GPU memory to serve the official FP8 weights (689 GB) with vLLM or SGLang, by our rule of weights × 1.2 for the KV cache and activations; about 485 GB with Ollama’s deepseek-r1:671b build. Long contexts and many parallel requests need more.
Can DeepSeek-R1-0528 run on a single GPU?
No. The smallest server that holds it is 2× MI355X, with Ollama + Open WebUI, deepseek-r1:671b (quantized), by our sizing rule.
What is the cheapest server for DeepSeek-R1-0528?
The cheapest CryptGPU server for DeepSeek-R1-0528 is 2× AMD MI355X at $2,818 a month, with Ollama + Open WebUI, deepseek-r1:671b (quantized). The model can be preinstalled at no extra cost when you order, and you pay in crypto with no KYC.
Does DeepSeek-R1-0528 run on the AMD MI355X?
Yes: its support on AMD Instinct GPUs with ROCm is confirmed for vLLM and SGLang, and the MI355X has 288 GB per GPU.
What licence does DeepSeek-R1-0528 use?
MIT. We download the official weights from Hugging Face (deepseek-ai/DeepSeek-R1-0528); you are responsible for the licence.
Other models to compare.
Your own DeepSeek-R1-0528, ready in minutes.
Dedicated GPUs, one monthly price, the model preinstalled. Paid in crypto, no KYC.
