GPU comparison · CDNA 4 vs Blackwell Ultra
AMD MI355X vs NVIDIA B300.
The two 288 GB GPUs. Both have 288 GB of HBM3E at 8 TB/s. The AMD MI355X adds strong FP64 and costs a fraction of the NVIDIA B300 per month; the B300 brings CUDA, more FP4 throughput and NVLink 5. Our MI355X market reference is indicative.
- 288 GB eachGPU memory
- 8 TB/s eachmemory bandwidth
- $1,409 vs $4,019per GPU, per month
MI355X vs B300: which one to rent.
Choose the MI355X if
- Your stack runs on ROCm: PyTorch, vLLM and SGLang support the MI355X.
- You want the most memory per dollar: $4.89 per GB per month against $13.95.
- You need FP64: 78.6 TFLOPS against 1.25 on the B300.
Choose the B300 if
- Your code depends on CUDA libraries, custom kernels or TensorRT-LLM.
- You run FP4 inference at scale: 13.5 PFLOPS dense per GPU against 10.1.
- You train across 8 GPUs and want NVLink 5 at 1.8 TB/s per GPU.
MI355X vs B300 specs, side by side.
Figures from NVIDIA’s and AMD’s published specifications, per GPU. Tensor figures are peak dense throughput.
| Per GPU | AMD MI355X | NVIDIA B300 | B300 / MI355X |
|---|---|---|---|
| Architecture | CDNA 4 (gfx950), TSMC 3 nm and 6 nm | Blackwell Ultra, TSMC 4NP | — |
| Form factor | OAM module, liquid-cooled | SXM module (HGX B300) | — |
| GPU memory | 288 GB HBM3E (8 stacks of 36 GB) | 288 GB HBM3E | Same |
| Memory bandwidth | 8 TB/s | Up to 8 TB/s | Same |
| GPU-to-GPU link | Infinity Fabric, 7 links of 153.6 GB/s | NVLink 5, 1.8 TB/s per GPU | 1.67× |
| FP4 tensor (dense) | 10.1 PFLOPS | 13.5 PFLOPS | 1.34× |
| FP8 tensor (dense) | 5,000 TFLOPS | 4,500 TFLOPS | 0.9× |
| FP16 / BF16 tensor (dense) | 2,500 TFLOPS | 2,250 TFLOPS | 0.9× |
| FP64 | 78.6 TFLOPS | 1.25 TFLOPS | 0.02× |
| Max power | 1,400 W | Not published | — |
| GPUs per server | 1, 2, 4, 8 | 1, 2, 4, 8 | — |
| Per GPU in our servers | 32 vCPU · 384 GB RAM · 3.84 TB | 32 vCPU · 384 GB RAM · 3.84 TB | — |
Sources: AMD Instinct MI355X · ROCm MI350 series · NVIDIA HGX · Blackwell Ultra technical blog. Full sheets: MI355X · B300
MI355X vs B300 rental price.
Our monthly prices, set 30% below the market median of public on-demand prices and paid in crypto. Hourly figures are the monthly price divided by 730 hours.
| CryptGPU | AMD MI355X | NVIDIA B300 | B300 / MI355X |
|---|---|---|---|
| Price per GPU, per month | $1,409 | $4,019 | 2.9× |
| Equivalent per GPU-hour | $1.93 | $5.51 | — |
| Per GB of GPU memory, per month | $4.89 | $13.95 | 2.9× |
| Per PFLOPS of dense FP8, per month | $282 | $893 | 3.2× |
| Largest server | 8× · $11,272/mo | 8× · $32,152/mo | — |
| Market median, per GPU | $2,022* | $5,745 | — |
| Below the median | −30% | −30% | — |
* Indicative market reference: no provider publishes an on-demand price for this GPU yet. Same price per GPU at every server size, no hourly metering. How we compare prices · Configure an MI355X server · Configure a B300 server
Which models fit on each.
Smallest server that holds each model, by precision. Rule of thumb with headroom for the KV cache, 92% of GPU memory usable.
| Model (total parameters) | MI355X · FP8 | B300 · FP8 | MI355X · 4-bit | B300 · 4-bit |
|---|---|---|---|---|
| Qwen3.5-9B | 1× | 1× | 1× | 1× |
| Gemma 4 31B | 1× | 1× | 1× | 1× |
| Llama 3.3 70B | 1× | 1× | 1× | 1× |
| gpt-oss-120b | 1× | 1× | 1× | 1× |
| DeepSeek-V4-Flash | 2× | 2× | 1× | 1× |
| Qwen3.5-397B-A17B | 2× | 2× | 1× | 1× |
| DeepSeek-R1 (671B) | 4× | 4× | 2× | 2× |
| Kimi K2.6 (1T) | 8× | 8× | 4× | 4× |
FP8 ≈ 1.2 bytes and 4-bit ≈ 0.65 bytes per parameter. — = larger than the biggest server of that GPU. Size another model · How much VRAM does an LLM need?
MI355X vs B300 FAQ.
More on each GPU: AMD MI355X · NVIDIA B300.
Why is the MI355X so much cheaper than the B300?
We price both 30% below the market median, and the markets differ: public MI355X offers start far below B300 on-demand prices, and ROCm software is younger than CUDA. Our MI355X reference is indicative, as no provider publishes an on-demand price yet.
Will my PyTorch code run on the MI355X?
Most PyTorch code runs unchanged with the ROCm build of PyTorch. Custom CUDA kernels need HIP ports. vLLM and SGLang provide ROCm builds.
Which is better for scientific computing?
The MI355X: 78.6 TFLOPS of FP64 against 1.25 on the B300, with the same 288 GB and 8 TB/s.
Rent the MI355X or the B300.
Dedicated servers with 1 to 8 GPUs, one monthly price, paid in crypto. Online in under 10 minutes, no KYC.
