New NVIDIA B300 · 288 GB HBM3e servers — from $4,019/mo per GPU

See the B300

Open model · Moonshot AI · Kimi K3 licence

Run Kimi K3
on your own GPUs.

Largest open-weight model we offer: multimodal, 1M context. 2.8T MoE · 104B active parameters, 1561 GB of official MXFP4 weights.

  • 1,873 GBGPU memory, MXFP4
  • 2.8T MoE · 104B activeparameters
  • 8× B300cheapest server, MXFP4
  • $32,152per month

How much VRAM Kimi K3 needs.

Weights × 1.2 for the KV cache and activations, with 92% of GPU memory usable: the rule our deploy page uses. Long contexts and many parallel requests need more.

GPU memory needed for Kimi K3 and cheapest server
Serving optionGPU memoryCheapest serverPer monthAction
vLLM or SGLang, official MXFP4 weights 1,873 GB 8× NVIDIA B300 $32,152/mo Deploy

Weights: moonshotai/Kimi-K3 · 1561 GB, MXFP4. How we estimate GPU memory

Which GPUs run Kimi K3.

Smallest server of each GPU that holds the model, and its monthly price.

Number of GPUs needed for Kimi K3 on each GPU model
GPUMXFP4Price, MXFP4
NVIDIA B300288 GB HBM3e $32,152/mo
NVIDIA B200180 GB HBM3e
AMD MI355X288 GB HBM3E n/a
NVIDIA H200141 GB HBM3e
NVIDIA H10080 GB HBM3
NVIDIA RTX PRO 600096 GB GDDR7 ECC
NVIDIA L40S48 GB GDDR6 ECC
NVIDIA RTX 509032 GB GDDR7
NVIDIA RTX 409024 GB GDDR6X
NVIDIA RTX 508016 GB GDDR7

n/a = ROCm support for this model is not confirmed with vLLM or SGLang. — = larger than the biggest server of that GPU. Size another model

Deploy Kimi K3 preinstalled.

Pick the model at the Software step of the deploy page: we install it with the engine you choose, at no extra cost.

  1. 1

    Pick the server

    The deploy page proposes the cheapest server that holds the model, and switches when you change the precision or the engine.

  2. 2

    Pick the engine

    vLLM or SGLang serve an OpenAI-compatible API.

  3. 3

    Pay and log in

    Pay the month in BTC, ETH, USDT, XMR or LTC, no KYC. Your server is online in under 10 minutes after confirmation, with root SSH access.

Guides: serving LLMs with vLLM · GPU servers for LLM inference · first steps on your server

Kimi K3 FAQ.

Every model: open models and their VRAM.

How much VRAM does Kimi K3 need?

About 1,873 GB of GPU memory to serve the official MXFP4 weights (1561 GB) with vLLM or SGLang, by our rule of weights × 1.2 for the KV cache and activations. Long contexts and many parallel requests need more.

Can Kimi K3 run on a single GPU?

No. The smallest server that holds it is 8× B300, with vLLM or SGLang, official MXFP4 weights, by our sizing rule.

What is the cheapest server for Kimi K3?

The cheapest CryptGPU server for Kimi K3 is 8× NVIDIA B300 at $32,152 a month, with vLLM or SGLang, official MXFP4 weights. The model can be preinstalled at no extra cost when you order, and you pay in crypto with no KYC.

Does Kimi K3 run on the AMD MI355X?

Its support on ROCm is not confirmed yet, so we offer it on NVIDIA servers.

What licence does Kimi K3 use?

Kimi K3 licence, with conditions or commercial restrictions: read the model card before you use it. We download the official weights from Hugging Face (moonshotai/Kimi-K3); you are responsible for the licence.

Your own Kimi K3, ready in minutes.

Dedicated GPUs, one monthly price, the model preinstalled. Paid in crypto, no KYC.

Customer area

Sign in to CryptGPU

Your orders, invoices and servers in one place.

New to CryptGPU?

Accounts are created at checkout: choose your server and software, then create your account in the Account step of your first order.

Deploy a server