AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Generalized Visual Language Models

Collected Oct 1, 2026

In a post, Lilian Weng focuses on one approach to vision language tasks: extending pre-trained generalized language models to consume visual signals. Traditionally, such systems use an object detection network as a vision encoder and a text decoder, and image captioning and visual question-answering have been studied for years. Weng groups vision language models (VLMs) into four buckets: translating images into embedding features that can be jointly trained with token embeddings; learning image embeddings that work as a prefix for a frozen pre-trained language model; using a cross-attention mechanism to fuse visual information into language model layers; and combining vision and language models without any training. She also lists decoding guided with vision-based scores and language as a communication interface.

The post describes several models. VisualBERT feeds text and image regions into BERT, using tokenized features, segmentation embeddings and position embeddings, trained on MS COCO with masked language modeling and sentence-image prediction. SimVLM uses a prefix language model with bidirectional attention on the prefix and causal attention on the main input, trained on ALIGN image-text pairs and C4 text. CM3 is an autoregressive model trained on close to 1T Web data, with images tokenized by VQVAE-GAN into 256 tokens. Frozen and ClipCap update only the vision module or a mapping network while keeping the language model frozen. VisualGPT uses a self-resurrecting encoder-decoder attention mechanism; VC-GPT adds cross-attention layers and a self-ensemble module. MERLOT trains on 6 million YouTube videos with transcribed speech. Flamingo connects a pretrained LM and CLIP vision encoder via a Perceiver resampler and gated cross-attention layers.

Weng also references datasets, evaluation tasks including visual question-answering, visual language reasoning, and video QA and understanding.

Read at Lilian Weng

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Processing images to generate text, such as image captioning and visual question-answering, has been studied for years. Traditionally such systems rely on an object detection network as a vision encoder to capture visual features and then produce text via a text decoder. Given a large amount of existing literature, in this post, I would like to only focus on one approach for solving vision language tasks, which is to extend pre-trained generalized language models to be capable of consuming visual signals .