TLDRocket
Sign in

Diffusion Models for Video Generation

Lilian Weng

Researchers are extending diffusion models from image synthesis to video generation, a more complex task requiring temporal consistency across frames. Video generation demands larger high-quality datasets and text-video pairs, which are harder to obtain than image datasets. This capability enables models to generate coherent multi-frame sequences rather than single static images.

Why it matters

Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. The task itself is a superset of the image case, since an image is a video of 1 frame, and it is much more challenging because: It has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model. In comparison to text or images, it is more difficult to collect large amounts of high-quality, high-dimensional video data, let along text-video pairs. 🥑 Required Pre-read: Please make sure you have read the previous blog on “What are Diffusion Models?” for image generation before continue here.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.