New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Open model · Mistral AI · Apache 2.0

Run Devstral Small 2
on your own GPUs.

Mistral’s coding-agent model for IDE and CLI agents. 24B parameters, 26 GB of official FP8 weights.

  • 31 GBGPU memory, FP8
  • 24Bparameters
  • 1× RTX 4090cheapest server, Ollama
  • $269per month

How much VRAM Devstral Small 2 needs.

Weights × 1.2 for the KV cache and activations, with 92% of GPU memory usable: the rule our deploy page uses. Long contexts and many parallel requests need more.

GPU memory needed for Devstral Small 2 and cheapest server
Serving optionGPU memoryCheapest serverPer monthAction
vLLM or SGLang, official FP8 weights 31 GB 2× NVIDIA RTX 4090 $538/mo Deploy
Ollama + Open WebUI, devstral-small-2:24b (quantized) 18 GB 1× NVIDIA RTX 4090 $269/mo Deploy

Weights: mistralai/Devstral-Small-2-24B-Instruct-2512 · 26 GB, FP8 · Ollama: devstral-small-2:24b, 15 GB. How we estimate GPU memory

Which GPUs run Devstral Small 2.

Smallest server of each GPU that holds the model, and its monthly price.

Number of GPUs needed for Devstral Small 2 on each GPU model
GPUFP8OllamaPrice, FP8
NVIDIA B300288 GB HBM3e $4,019/mo
NVIDIA B200180 GB HBM3e $3,439/mo
AMD MI355X288 GB HBM3E n/a
NVIDIA H200141 GB HBM3e $2,279/mo
NVIDIA H10080 GB HBM3 $1,779/mo
NVIDIA RTX PRO 600096 GB GDDR7 ECC $959/mo
NVIDIA L40S48 GB GDDR6 ECC $789/mo
NVIDIA RTX 509032 GB GDDR7 $658/mo
NVIDIA RTX 409024 GB GDDR6X $538/mo
NVIDIA RTX 508016 GB GDDR7 $796/mo

n/a = ROCm support for this model is not confirmed with vLLM or SGLang. — = larger than the biggest server of that GPU. Size another model

Deploy Devstral Small 2 preinstalled.

Pick the model at the Software step of the deploy page: we install it with the engine you choose, at no extra cost.

  1. 1

    Pick the server

    The deploy page proposes the cheapest server that holds the model, and switches when you change the precision or the engine.

  2. 2

    Pick the engine

    vLLM or SGLang serve an OpenAI-compatible API; Ollama with Open WebUI gives a private chat in the browser.

  3. 3

    Pay and log in

    Pay the month in BTC, ETH, USDT, XMR or LTC, no KYC. Your server is online in under 10 minutes after confirmation, with root SSH access.

Guides: serving LLMs with vLLM · GPU servers for LLM inference · first steps on your server

Devstral Small 2 FAQ.

Every model: open models and their VRAM.

How much VRAM does Devstral Small 2 need?

About 31 GB of GPU memory to serve the official FP8 weights (26 GB) with vLLM or SGLang, by our rule of weights × 1.2 for the KV cache and activations; about 18 GB with Ollama’s devstral-small-2:24b build. Long contexts and many parallel requests need more.

Can Devstral Small 2 run on a single GPU?

Yes, on one B300, B200, MI355X, H200, H100, RTX PRO 6000, L40S, RTX 5090 or RTX 4090 with Ollama + Open WebUI, devstral-small-2:24b (quantized), by our sizing rule.

What is the cheapest server for Devstral Small 2?

The cheapest CryptGPU server for Devstral Small 2 is 1× NVIDIA RTX 4090 at $269 a month, with Ollama + Open WebUI, devstral-small-2:24b (quantized). The model can be preinstalled at no extra cost when you order, and you pay in crypto with no KYC.

Does Devstral Small 2 run on the AMD MI355X?

Not confirmed with vLLM or SGLang on ROCm yet, so we offer it on NVIDIA servers; with Ollama it also runs on the MI355X.

What licence does Devstral Small 2 use?

Apache 2.0. We download the official weights from Hugging Face (mistralai/Devstral-Small-2-24B-Instruct-2512); you are responsible for the licence.

Your own Devstral Small 2, ready in minutes.

Dedicated GPUs, one monthly price, the model preinstalled. Paid in crypto, no KYC.

Customer area

Sign in to CryptGPU

Your orders, invoices and servers in one place.

New to CryptGPU?

Accounts are created at checkout: choose your server and software, then create your account in the Account step of your first order.

Deploy a server