Amazon launched the SageMaker HyperPod Inference Gateway, a Kubernetes-native routing addon for LLM inference on EKS
Feature update Provisional 86% confidence first seen
Amazon introduced the SageMaker HyperPod Inference Gateway, a GPU-aware Kubernetes routing component for running LLM inference on Amazon EKS with SageMaker HyperPod. The company said it reduces first-token latency by steering inference requests to the best-suited model pods using real-time GPU signals and routing rules, and reported benchmark improvements for Llama-3.1-70B under bursty traffic. A separate SageMaker AI post reviewed 2026 year-to-date launches related to its managed endpoints and HyperPod inference deployment paths.