NeoMME: an efficient Multimodal-native and Multilingual Encoder
Hugging Face
Read full postNeoMME is a new family of multilingual multimodal encoders with 260M and 800M parameters that use a single bidirectional Transformer to process text and raw image patches without separate pretrained vision or language models. It achieves efficient visual document retrieval, outperforming previous models in speed and storage efficiency, and is available on Hugging Face under Apache 2.0 license.


