Tencent Releases WeMM-Embedding Models, Claiming Top Scores on Multimodal Benchmarks
Summary
Tencent's WeChat Vision Team releases WeMM-Embedding, a family of multimodal AI models in 2B, 4B, and 9B sizes that claim top scores on MMEB-v2 and MMEB-v3 benchmarks, supporting text, images, videos, and documents, now publicly available on Hugging Face under Apache License 2.0.
Key Points
- Tencent's WeChat Vision Team releases WeMM-Embedding, a family of universal multimodal embedding models available in 2B, 4B, and 9B parameter sizes, supporting text, images, videos, visual documents, and interleaved multimodal inputs.
- The models achieve state-of-the-art performance on MMEB-v2 and MMEB-v3 benchmarks, with the 9B model scoring 80.6 on MMEB-v2 and 59.5 on MMEB-v3, outperforming competitors including Qwen3-VL-Embedding and GME across image, video, and visual document tasks.
- The models support Matryoshka Representation Learning for flexible embedding dimensions, are compatible with Transformers, Sentence Transformers, vLLM, and SGLang, and are publicly available on Hugging Face under the Apache License 2.0.