The Tencent WeChat Vision Team has officially released the open-source universal multimodal embedding model, WeMM-Embedding. This innovative model is designed to accommodate multimodal inputs, including text and images, and is available in three distinct versions: 2B, 4B, and 9B. Leveraging a unified architecture, two-stage training process, and knowledge distillation techniques, WeMM-Embedding has been successfully deployed on a large scale within WeChat's recommendation and search systems. Impressively, it handles over one billion daily online queries, providing robust technical support for AI application developers and fostering the development of innovative AI applications across various vertical domains.
