Diffusion Models for Video Generation
Diffusion models have shown strong results on image synthesis, and the research community has started applying them to the harder task of video generation, according to a blog post by Lilian Weng. The post notes the task is a superset of image generation, since an image is a one-frame video, and is more challenging due to requirements for temporal consistency across frames and the difficulty of collecting large amounts of high-quality video data, including text-video pairs.
The post reviews approaches for designing and training diffusion video models from scratch, without relying on pre-trained image generators. It covers parameterization and sampling basics, including a v-prediction parameterization proposed by Salimans and Ho (2022), which has been shown to help avoid color shift in video generation compared to epsilon-parameterization. It also describes reconstruction guidance from Video Diffusion Models (VDM; Ho and Salimans, et al. 2022) for conditioning the sampling of one video segment on another, such as autoregressive extension or filling in missing frames.
On architecture, the post states that U-Net and Transformer remain common choices. VDM extends the 2D U-Net to 3D data, factorizing layers over space and time, with a temporal attention block added after each spatial attention block to capture temporal coherence. Imagen Video (Ho, et al. 2022) uses a cascade of seven diffusion models, including a frozen T5 text encoder, a base video model, and three temporal and three spatial super-resolution models, outputting 1280x768 videos at 24 fps. Sora (Brooks et al. 2024) uses a Diffusion Transformer architecture operating on spacetime patches of video and image latent codes.
For adapting image models, the post describes inflating a pre-trained text-to-image diffusion model by inserting temporal layers. Make-A-Video (Singer et al. 2022) extends a pre-trained image model with spatiotemporal convolution and attention layers and a frame interpolation network. Tune-A-Video (Wu et al. 2023) enables one-shot video tuning for object editing, background changes and style transfer using a spatiotemporal attention block.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. The task itself is a superset of the image case, since an image is a video of 1 frame, and it is much more challenging because: It has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model. In comparison to text or images, it is more difficult to co