AI flagship · CDNA 4
AMD MI355X servers Best value
AMD’s CDNA 4 accelerator: 288 GB of HBM3E at a fraction of NVIDIA flagship prices.
- 288 GBHBM3E per GPU
- 8 TB/smemory bandwidth
- Infinity FabricGPU interconnect
- 1–8 GPUsper server
Price per GPU · monthly
$1,409/mo
−30%Market median $2,022≈ $1.93 per GPU-hour
Below every on-demand list price we found
Server size
Configure a MI355X serverUSD per server per month · prepaid · pay in BTC, ETH, USDT, USDC, LTC
Why the MI355X.
The Instinct MI355X is AMD’s latest data-center GPU available to rent. It carries 288 GB of HBM3E at 8 TB/s, supports FP4, FP6 and FP8, and keeps strong FP64. Its public price is low because on-demand supply is limited and the software ecosystem is ROCm rather than CUDA: if your stack runs on ROCm, it is the most memory per dollar in our range.
- 288 GB HBM3E
Hold a 405B model in FP8 on two GPUs by our sizing rule.
- Most memory per dollar
About $4.89 per GB per month, against $13.95 for the B300.
- Strong FP64
Suited to scientific computing as well as AI.
From 1 to 8 MI355X GPUs.
Same price per GPU at every size. vCPU, DDR5 RAM and NVMe scale with the number of GPUs.
| Server | GPU memory | vCPU | RAM | NVMe | Market median | CryptGPU | Action |
|---|---|---|---|---|---|---|---|
| 1× MI355X | 288 GB | 32 | 384 GB | 3.84 TB | $1,409/mo | Configure 1× MI355X | |
| 2× MI355XInfinity Fabric | 576 GB | 64 | 768 GB | 7.68 TB | $2,818/mo | Configure 2× MI355X | |
| 4× MI355XInfinity Fabric | 1,152 GB | 128 | 1,536 GB | 15.36 TB | $5,636/mo | Configure 4× MI355X | |
| 8× MI355XInfinity Fabric | 2,304 GB | 256 | 3,072 GB | 30.72 TB | $11,272/mo | Configure 8× MI355X |
AMD MI355X specifications.
Figures from AMD’s published specifications. Tensor figures are peak dense throughput unless marked otherwise.
GPU and memory
- Architecture
- CDNA 4 (gfx950), TSMC 3 nm and 6 nm
- Form factor
- OAM module, liquid-cooled
- GPU memory
- 288 GB HBM3E (8 stacks of 36 GB)
- Memory bandwidth
- 8 TB/s
- GPU-to-GPU link
- Infinity Fabric, 7 links of 153.6 GB/s
- Host link
- PCIe 5.0 x16
- Board power
- 1,400 W
- Partitioning
- 1, 2, 4 or 8 compute partitions
Matrix throughput (dense)
- FP4 / FP6
- 10.1 PFLOPS
- FP8
- 5 PFLOPS
- FP16 / BF16
- 2.5 PFLOPS
- INT8
- 5 POPS
- FP32
- 157.3 TFLOPS
- FP64
- 78.6 TFLOPS
Compute units
- Compute units
- 256
- Stream processors
- 16,384
- Matrix cores
- 1,024
Server, per GPU
- vCPU
- 32
- RAM
- 384 GB DDR5
- NVMe
- 3.84 TB
- GPUs per server
- 1, 2, 4, 8
AMD figures, dense. AMD lists no TF32 rate: ROCm emulates TF32 through BF16. Partitioning replaces NVIDIA’s Multi-Instance GPU on AMD Instinct.
Which models fit on MI355X servers.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache; long contexts and large batches need more.
| Model (total parameters) | FP16 | FP8 | 4-bit |
|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× |
| Llama 3.3 70B | 1× | 1× | 1× |
| gpt-oss-120b | 2× | 1× | 1× |
| DeepSeek-V4-Flash | 4× | 2× | 1× |
| Qwen3.5-397B-A17B | 4× | 2× | 1× |
| DeepSeek-R1 (671B) | 8× | 4× | 2× |
| Kimi K2.6 (1T) | — | 8× | 4× |
Estimate: FP16 ≈ 2.4 bytes, FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter, with 92% of GPU memory usable. — = does not fit in the largest MI355X server. Several recent models ship natively in FP8 or 4-bit formats; the columns show what each precision needs. Try the sizing helper for other models and for fine-tuning.
MI355X prices across the market.
Monthly price per GPU. Public on-demand list prices collected on 23 Sep 2026; hyperscalers shown for reference only.
Per GPU, 730 hours a month. No provider publishes an on-demand price for this GPU yet: its market reference is the median of the public prices we found. Hyperscalers (AWS, Google Cloud, Azure, Oracle) are excluded from the median. How we compare
MI355X questions.
More in the full FAQ.
Will my PyTorch code run on the MI355X?
Most PyTorch code runs unchanged with the ROCm build of PyTorch. Custom CUDA kernels need HIP ports; check the libraries you depend on.
Why is the MI355X so much cheaper than the B300?
Both have 288 GB, but the MI355X is priced from a much lower market: the public offers we found are starting prices far below B300 on-demand prices, and the ROCm software ecosystem is younger than CUDA. We price 30% below the median of those public prices.
Can I run vLLM on it?
Yes, vLLM and SGLang both provide ROCm builds for AMD Instinct GPUs.
Build your MI355X server.
Pick 1 to 8 GPUs and see the exact monthly price. $1,409 per GPU, the same at every size.
