AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Multimodality and Large Multimodal Models (LMMs)

Collected Oct 1, 2026

Chip Huyen published a post on multimodality and Large Multimodal Models (LMMs), noting that machine learning models historically handled a single data mode, such as text, image, or audio. The post cites OpenAI's GPT-4V system card as saying that incorporating additional modalities, such as image inputs, into LLMs is viewed by some as a key frontier in AI research and development.

According to the post, adding modalities to Large Language Models produces LMMs, but not all multimodal systems are LMMs: text-to-image models like Midjourney, Stable Diffusion, and Dall-E are multimodal yet lack a language model component. Multimodal can describe differing input and output modalities, multimodal inputs, or multimodal outputs.

The post has three parts. Part 1 covers context for multimodality, data modalities, and task types, splitting vision-language tasks into generation and vision-language understanding. Part 2 covers fundamentals using CLIP, which Huyen says was the first model to generalize to multiple image classification tasks with zero- and few-shot learning, and Flamingo, whose strong performance prompted some to consider it the GPT-3 moment in the multimodal domain. The post states Flamingo was not the first large multimodal model that could generate open-ended responses, as Salesforce's BLIP came out three months prior.

Part 3 discusses research areas including incorporating more data modalities, instruction-following multimodal systems, adapters for more efficient multimodal training, and generating multimodal outputs, covering BLIP-2, LLaVA, LLaMA-Adapter V2, and LAVIN. Huyen describes CLIP's shared embedding space and notes Flamingo and LLaVa use CLIP as their image encoder, DALL-E uses it to rerank generated images, and it is unclear whether GPT-4V uses CLIP.

Read at Chip Huyen

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

For a long time, each ML model operated in one data mode – text (translation, language modeling), image (object detection, image classification), or audio (speech recognition). However, natural intelligence is not limited to just a single modality. Humans can read, talk, and see. We listen to music to relax and watch out for strange noises to detect danger. Being able to work with multimodal data is essential for us or any AI to operate in the real world. OpenAI noted in their GPT-4V system card that “ incorporatin