Performance per dollar is getting faster and cheaper
TLDR Dev
AMD's MI355X GPU achieves comparable inference performance to NVIDIA's Blackwell at roughly 2.75x lower cost per unit, though it historically lagged due to software support issues. Wafer demonstrated 2626 tokens per second aggregate throughput on the MI355X versus 3192 on Blackwell, and 213 tokens per second on GLM5.2, by optimizing quantization, speculative decoding, and kernel selection. As software tooling and agent-driven optimization improve, AMD's cost advantage is becoming increasingly accessible without requiring custom kernel development.
Why it matters
The release of GLM-5.2 shows advancements in inference performance, with 2626 tok/s/node on AMD MI355X at more than twice the cost-efficiency compared to similar NVIDIA setups. Despite historical challenges with AMD's software support, improvements in model optimization are narrowing the gap.