Evolution of the PyTorch Media Processing Landscape

PyTorch has consolidated its media processing stack over the last two years, moving all decoding and encoding of images, video, and audio into a single library, TorchCodec. The equivalent APIs that previously lived in TorchVision and TorchAudio are now deprecated or removed. TorchVision and TorchAudio are instead focused on transforms, with models, datasets, and pipelines no longer under active development.
Previously, decoding and encoding were scattered and partly duplicated. TorchVision offered io.read_video() and io.VideoReader() across three backends: PyAV in Python, a C++ FFmpeg backend, and a CUDA/NVCUVID backend, though only PyAV worked out of the box since the others required building from source and were tied to a specific FFmpeg version. TorchAudio had StreamReader and StreamWriter plus audio decoding utilities built on FFmpeg, libsoundfile, and libsox. Image decoding lived in TorchVision. TorchCodec now handles all of this on CPU and CUDA.
PyTorch gives three reasons: a single location so users do not have to guess which library has a given decoder; concentrated optimization effort, with TorchCodec generally more performant than the old TorchVision and TorchAudio implementations, particularly for CUDA video decoding; and simpler maintenance, since media I/O pulls in six major FFmpeg versions (4 through 9 at the time of writing), NVIDIA's codec SDK, and one C/C++ library per image format such as libjpeg and libpng, each with its own licensing rules.
The division of labor is direct. TorchCodec turns media files or encoded bytes into tensors and tensors back into files, while TorchVision and TorchAudio transform the tensors in between. Video uses TorchCodec's VideoDecoder with TorchVision v2 transforms; audio uses TorchCodec's AudioDecoder with TorchAudio transforms; images use decode_image and feed torchvision.transforms.v2. Encoding uses VideoEncoder, AudioEncoder, JpegEncoder, and PngEncoder. TorchAudio's transition was particularly disruptive, with many APIs deprecated and then removed, though community feedback led PyTorch to keep several popular APIs originally slated for removal.
All three libraries are now ABI stable, so a given version is not tied to a single PyTorch version and keeps working with later PyTorch releases. They no longer need rebuilding for each PyTorch release, and their release cadence no longer matches PyTorch's.
Why it matters: developers still using decoding or encoding APIs in TorchVision or TorchAudio should migrate to TorchCodec, which PyTorch points to via a migration guide, and can report missing functionality on GitHub. Existing pinned versions of the transformed libraries continue to work with newer PyTorch builds.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
TL;DR If you need to decode or encode media, whether it’s images, video, or audio, use TorchCodec. If you need to transform media, use TorchVision for images and video, and... The post Evolution of the PyTorch Media Processing Landscape appeared first on PyTorch .