Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM
Amazon Web Services Dmitry Soldatkin
Alibaba’s Qwen team released Qwen3.8-2.4T-A95B as open weights and then documented how to run it on Amazon SageMaker HyperPod with vLLM. The deployment uses a ml.p6-b300.48xlarge instance with 8× NVIDIA B300 Blackwell Ultra GPUs, and it applies NVFP4 quantization to fit the model’s ~1.2 TB weights on the node. As a result, teams can serve a 2.4T-parameter Qwen-Max-class model through an OpenAI-compatible endpoint with controlled reasoning, tool calling, and native MTP speculative decoding while avoiding per-token API fees at scale.
Why it matters
Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.
Related stories
[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
Latent Space · 1 month ago ·
2
Introducing Qwen1.5
GitHub Pages · 2 years ago ·
20
Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and the Most Capable One in the Qwen Family to Date
MarkTechPost · 1 month ago ·
48