AI article

NeoMME: an efficient Multimodal-native and Multilingual Encoder

No preview is available. Read the original article for the full story.

huggingface | Sep 3, 2026 | Tony Wu, Aurélien Lac

Automated excerpt

We fine-tuned NeoMME for visual document retrieval using ColPali's page-image approach. NeoMME encoder backbone One Transformer for images and text NeoMME comes in two sizes, 260M and 800M. The image patches remain visible while NeoMME reconstructs masked text. Pretraining mixes multilingual text, code, mathematics, natural images, and document images. Late-interaction and dense retrieval heads for both NeoMME model sizes.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news