AivexaNewsSearch
AI news for builders and product teamsChecked every hour

NeoMME: an efficient Multimodal-native and Multilingual Encoder

Collected Oct 1, 2026

Hugging Face introduced NeoMME, a family of 260M and 800M multilingual multimodal encoders. Unlike many generative visual language models, NeoMME does not use a separate pretrained vision tower or a causal language model. A single bidirectional Transformer processes both text tokens and non-overlapping 32x32 image patches, and the entire model is trained from scratch with a masked discrete-diffusion objective.

NeoMME supports dynamic image resolution, a context length of 16,384 tokens, and a BPE tokenizer with a 131k-token vocabulary built on multilingual text, code, mathematics, and machine-produced image transcripts. Each model processes about 524 billion packed input tokens, including 290 billion text-only tokens; the NorMuon optimizer was chosen for data efficiency.

NeoMME-Retriever is fine-tuned for visual document retrieval following ColPali's page-image approach, with a dense head (mean pooling) and a late-interaction head projecting tokens and patches to 128-dimensional normalized vectors. One forward pass returns both representations. On ViDoRe v3, NeoMME-Retriever-260M reaches 0.523 nDCG@10, described as the highest score among evaluated models strictly below 800M parameters and within 0.002 of ColQwen2.5 with about 14x fewer parameters. The 800M model reaches 0.556, within 0.009 of the similarly sized Vultron Retriever Flash (0.8B). Both are said to lie on the model-size Pareto frontier.

At a matched 2048x2048 input size on one NVIDIA L40S GPU, the 260M model encodes about 51 pages per second, roughly twice ColModernVBERT's 26 pages per second. Hierarchical token pooling and asymmetric quantization reduce late-interaction index storage; one configuration cuts storage from about 1.5 MB to 39 kB per page while retaining over 99% of baseline nDCG@10, and a more aggressive configuration reaches 6 kB per page (255x smaller) while retaining more than 95%.

NeoMME is available in Hugging Face Transformers, and all model checkpoints are released under the Apache 2.0 license. A technical report, model collection, and a visual RAG demo are also provided. The citation lists authors Aurélien Lac and Tony Wu, with H Company acknowledged for support and compute.

Read at Hugging Face Blog

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt