WeChat Vision · Tencent

WeMM-Embedding

WeChat Multi-Modal Embedding

  • Junjie Zhou*
  • Ke Mei*†
  • Lei Li*
  • Tianyi Wang
  • Fengyun Rao
  • Jing Lyu

WeChat Vision, Tencent Inc.* Core contributors† Project leader‡ Corresponding author

Universal representations for text, images, videos, visual documents, and mixed inputs. Built for multimodal understanding, retrieval, recommendation, and beyond.

Open model family2B4B9B

MULTIMODAL REPRESENTATIONS

Unified Embedding Space

Different inputs. A shared representation.

Multimodal inputs become embeddings Text, images, video, and visual documents pass through WeMM-Embedding. Four aligned rows show their vectors with the same dimension structure in a shared representation space. Cell intensities are conceptual, not measured embedding values or similarity scores. TEXT A turtle on amossy log. IMAGE VIDEO DOCUMENT WeMM-Embedding General Multimodal Embedding Models SHARED EMBEDDING SPACE COMMON DIMENSIONS → TEXTIMAGEVIDEODOCUMENT Different modalities. The same representation format. Representation concept · illustrative representations

01 / Method overview

From broad alignment
to fine-grained relevance.

Two-stage training turns a native multimodal foundation into general-purpose embeddings.

Read the methodology
STAGE 1 · MULTIMODAL ALIGNMENT

Learn a shared representation space.

Large-scale paired data across modalities, tasks, and domains establish broad coverage in a shared embedding space.

Semantic-ID-guided resampling balances training data Four illustrative semantic groups have different frequencies in the source data. Semantic IDs guide density-aware resampling: frequent groups are sampled at lower rates, while rarer groups are retained at higher rates. The curated distribution remains nonuniform. Bar lengths are conceptual, not measured dataset counts. SOURCE DATA CURATED DATA A A B B C C D D SID Resampling Less repetition. Broader semantic coverage. SEMANTIC GROUPS A–D · ILLUSTRATIVE DISTRIBUTIONS
STAGE 2 · DATA CURATION

Balance the training data.

Semantic IDs guide resampling to reduce overrepresented patterns and improve semantic balance. Hard negatives further strengthen fine-grained discrimination.

STAGE 2 · KNOWLEDGE TRANSFER

Learn the teacher’s similarity distribution.

For the same query and candidate samples, the 2B and 4B students learn to match the 9B teacher’s similarity distribution.

02 / Applications

Deployed across
WeChat applications.

  • Moments

  • Channels

  • Official Accounts

  • Weixin Shop

WeMM-Embedding powers search and recommendation at scale across WeChat Channels, Official Accounts, Moments, and e-commerce services. It delivers consistent gains across 14 online A/B tests.

03 / Benchmarks

Performance
across modalities.

Results across image, video, visual-document and cross-modal retrieval, alongside in-house evaluations.

Technical-report figure: WeMM-Embedding achieves MMEB-v2 scores of 77.9 at 2B, 79.2 at 4B and 80.6 at 9B, with comparisons across image, video, cross-modal retrieval and in-house evaluations.
Figure 1 · Public benchmarks and in-house evaluations.Original figure · PDF ↗
MMEB-v2 · Full results & evaluation details

Source: official repository · Evaluation code ↗

04 / Selected benchmark cases

Retrieval in practice.

Six real retrieval examples from WeMM-Embedding-9B. Inspect the correct Top-1 match alongside the next two candidates.

BENCHMARK EXPLORERWeMM-Embedding-9B

Loading benchmark examples…

Selected successful cases · recorded results · no live inference

Benchmark media is shown in limited form for research and evaluation illustration. Copyright remains with the respective owners. Media credits ↗

05 / Open models

Try the models.

Code & documentation

All three models support text, images, video, visual documents, and mixed inputs, with flexible embedding dimensions.

Compact representationsExplore Matryoshka dimensions

WeMM-Embedding-2B · MRL

Smaller vectors.
Measured performance.

One model. Multiple embedding dimensions.
Choose the representation size for your application.

WeMM-Embedding-2B · MMEB-v2
Relative to the full 2,048-dimensional representation.

View the reported analysis
256
Image performance retained
Video performance retained
Raw vector storage

For 1 million vectors, FP16, decimal GB. Excludes index overhead, metadata, and model weights. Retention values apply to image/video tasks, not visual documents.

Inspect the retrieved image