TLDRocket
Sign in

Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod

Amazon Web Services Nathan Arnold

NVIDIA and AWS describe how to run a continuous Physical AI “model factory” pipeline using NVIDIA Cosmos 3 on Amazon SageMaker HyperPod, moving through synthetic-data generation, post-training, and closed-loop evaluation on a single shared cluster and storage setup. Cosmos 3-Super is a 64B-parameter model that can be post-trained into a deployable policy. The workload no longer requires separate GPU pools per pipeline stage because the same persistent node pool under one control plane time-shares generation, training, and evaluation for better GPU goodput across the whole loop.

Why it matters

Building a Physical AI system takes a continuous pipeline, not a single training job. This post shows how to run that model factory (synthetic data generation, post-training, and closed-loop evaluation with NVIDIA Cosmos 3) on a persistent, resilient Amazon SageMaker HyperPod cluster on Amazon EKS, with GPU goodput as the metric that matters.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.