TLDRocket
Sign in

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

Amazon Web Services Dmitry Soldatkin

Alibaba’s Qwen team released Qwen3.8-2.4T-A95B as open weights and then documented how to run it on Amazon SageMaker HyperPod with vLLM. The deployment uses a ml.p6-b300.48xlarge instance with 8× NVIDIA B300 Blackwell Ultra GPUs, and it applies NVFP4 quantization to fit the model’s ~1.2 TB weights on the node. As a result, teams can serve a 2.4T-parameter Qwen-Max-class model through an OpenAI-compatible endpoint with controlled reasoning, tool calling, and native MTP speculative decoding while avoiding per-token API fees at scale.

Why it matters

Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.