Generate images and video with vLLM-Omni on SageMaker AI – Part 2
Amazon Web Services Yadan Wei ● Covered by 2 sources
AWS shows how to turn one prompt into an image, then a short video, on SageMaker AI. The image uses real-time inference; the video waits its turn and lands in S3.
Based on reporting by Amazon Web Services, Yadan Wei — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
AWS has split a generative media demo into two separate moves: first, a text prompt becomes an image; then that image gets animated into a short video on Amazon SageMaker AI. The setup uses the same AWS vLLM-Omni Deep Learning Container for both jobs, but two different endpoint styles. FLUX.2-klein-4B handles image generation on a real-time endpoint. Wan2.1-VACE-1.3B handles video generation on an asynchronous endpoint.
That split is the point. The image endpoint returns a base64 PNG inline, which the application can use immediately. The video endpoint takes a multipart request that includes the generated image and a motion prompt, then writes the MP4 to Amazon S3 when it’s done. AWS says this makes the different payloads and response patterns easier to see than in the earlier speech-focused Part 1 of the series.
The container story matters too. The post uses a pinned vLLM-Omni DLC release, called omni-sagemaker-cuda-v1.6, and changes the SM_VLLM_MODEL setting so each endpoint loads a different model. AWS says keeping one container image reduces serving-stack variation, while separate endpoints let each model use the instance type and inference option that fits its workload. The sample uses ml.g6.xlarge for the image endpoint and ml.g6e.xlarge for the video endpoint.
The code path is pretty direct. A prompt such as “Cinematic photograph of a coastal observatory at sunrise” goes to FLUX.2-klein. The app then resizes the returned image to 480 by 320, converts it into a JPEG data URL, and passes that into the Wan VACE request with 17 frames and 30 diffusion steps. AWS says the multipart body is uploaded to S3 before asynchronous inference reads it back, and the resulting MP4 is saved under outputs/.
There’s also a Streamlit front end for people who prefer clicking over terminals. And there are a few practical notes buried in the walkthrough: the video path can use Inline Body only up to 128,000 bytes, so this sample uses InputLocation instead; four diffusion steps are only for a smoke test; and the cleanup script leaves generated objects in S3 until someone removes them. The post is a good reminder that “one model, one endpoint” is a lazy default, not a law.
My take — AI-written commentary, not fact-checked reporting
This is the sensible kind of AI infrastructure post: less sparkle, more plumbing. Open ecosystems win when they can mix models, payloads, and serving patterns without pretending every workload deserves the same latency fantasy. The real headline here is not image-to-video magic; it’s AWS admitting that one size does not fit all, which is rarer than it should be.
Read more about this at: Amazon Web Services