Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
Amazon Web Services Nathan Arnold
NVIDIA and AWS describe how to run a continuous Physical AI “model factory” pipeline using NVIDIA Cosmos 3 on Amazon SageMaker HyperPod, moving through synthetic-data generation, post-training, and closed-loop evaluation on a single shared cluster and storage setup. Cosmos 3-Super is a 64B-parameter model that can be post-trained into a deployable policy. The workload no longer requires separate GPU pools per pipeline stage because the same persistent node pool under one control plane time-shares generation, training, and evaluation for better GPU goodput across the whole loop.
Why it matters
Building a Physical AI system takes a continuous pipeline, not a single training job. This post shows how to run that model factory (synthetic data generation, post-training, and closed-loop evaluation with NVIDIA Cosmos 3) on a persistent, resilient Amazon SageMaker HyperPod cluster on Amazon EKS, with GPU goodput as the metric that matters.
Related stories
Into the Omniverse: How Open World Models Push the Frontier of Physical AI
NVIDIA · 4 weeks ago ·
17
Deploying Kimi K3 on Amazon SageMaker HyperPod and Amazon EKS
AWS · 1 month ago ·
45
NVIDIA's GTC 2025 Announcement for Physical AI Developers: New Open Models and Datasets
Hugging Face · 1 year ago ·
53