AivexaNewsSearch
AI news for builders and product teamsChecked every hour

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Collected Oct 6, 2026

Google DeepMind launched EmbeddingGemma 2, described as the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space. The model was announced by Sahil Dua and Henrique Schechter Vera, both research engineers at Google DeepMind.

EmbeddingGemma 2 is built on the Gemma 4 architecture and released under a commercially permissive Apache 2.0 license. It has 740 million parameters and is positioned for on-device inference. It requires as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support.

The model features an 8K token context window, four times larger than EmbeddingGemma 1, and can process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations on local hardware. Using Matryoshka Representation Learning, developers can truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions, providing up to 6x storage reduction for local vector databases.

According to Google DeepMind, EmbeddingGemma 2 achieves leading scores among sub-1B multimodal embedders across benchmarks including MTEB Code and MAEB, while matching or outperforming many larger models across text, vision, and audio tasks. It records a 9.92-point improvement on code performance in MTEB Code, from 68.76 to 78.68.

With quantization on a Google Pixel 11 Pro, the model requires as little as roughly 191MB active RAM for text-only weights and roughly 567MB for the full multimodal model. It shares the text tokenizer and audio encoder with Gemma 4, allowing both to run together in a unified pipeline with a lower combined total memory footprint.

EmbeddingGemma 1, introduced last year for text embeddings, surpassed 20 million downloads, Google DeepMind said. Model weights for EmbeddingGemma 2 are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform Model Garden availability coming soon. Supported tooling includes transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, with vector storage via Qdrant.

Read at Google DeepMind

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt