AI article
NeoMME: an efficient Multimodal-native and Multilingual Encoder
No preview is available. Read the original article for the full story.
huggingface | Sep 3, 2026 | Tony Wu, Aurélien Lac
Automated excerpt
We fine-tuned NeoMME for visual document retrieval using ColPali's page-image approach. NeoMME encoder backbone One Transformer for images and text NeoMME comes in two sizes, 260M and 800M. The image patches remain visible while NeoMME reconstructs masked text. Pretraining mixes multilingual text, code, mathematics, natural images, and document images. Late-interaction and dense retrieval heads for both NeoMME model sizes.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.