A kernel-centric path to real-time video generation on Trainium
The AWS Neuron Science team and Reactor, a platform for deploying and scaling real-time interactive AI models, collaborated to enable real-time video generation on Trainium, Amazon's AI chip. The work was presented as a kernel-centric optimization effort using the Neuron Kernel Interface (NKI).
Jun Wu, a principal applied scientist on the Neuron team, said the team explores techniques for generative AI model enablement optimization, including new architectures, algorithms, and ways to generate and optimize code for running models on Trainium. Wu said the team noticed an evolution in video generation from shorter, fixed-length videos to infinite-length or dynamic-length videos, which gave rise to autoregressive diffusion models that combine next-token prediction with diffusion iteration.
Reactor co-founder and CTO Bryce Schmidtchen said Reactor had been working on real-time and interactive video generation for over 18 months. The teams aligned on Rolling Forcing, citing its ability to consistently generate high-quality 30-second videos and its relative size. Wu said Rolling Forcing is relatively small but has a very large sequence length, and that its 16 frames per second matched video playback and latency requirements.
The teams focused on three challenges for generic compilers: dynamic shapes, unusual memory access patterns, and heavy cache management. Using NKI, Wu said the 3D-RoPE kernel went from five seconds to 1.8 milliseconds, cache copies from 23 milliseconds to 1.9 milliseconds per layer, and attention transposes were eliminated by fusing them into the attention kernel. Wu also said NKI-Dev-Suite, an agent for generating NKI kernels, produced a working 3D-RoPE kernel on its first attempt, and that the pipeline used 11 GB of high-bandwidth memory while the standard eager-mode path ran out of memory.
Lingfan Yu, a senior applied scientist, said self-attention operates on 23,400 query tokens attending to 32,760 context tokens and accounts for about 70% of compute time. The team used hybrid sharding combining sequence parallelism and tensor parallelism, splitting heads across 4 cores and sequences across 2. Yu said the VAE decoder used spatial W-axis sharding, achieving a super-linear 8.25 times speedup.
The teams also batched common components of the diffusion and cache update phases. After these and other optimizations, they said they generated a correct video on the first end-to-end run. Yahav Biran, a principal solutions architect, said Rolling Forcing is a small model but a robust end-to-end system with an encoder, DIT, VAE, and decoder. Yu said the team is building common techniques for models employing autoregressive diffusion that require real-time interaction, not optimizing a single model only.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Using the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.