Skip to content
Preprint

Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

A disentangled contrastive learning~(DCL) method for multilingual dense retrieval by separating multilingual representations into semantic and linguistic subspaces based on hierarchical semantic alignment and language debiasing contrastive learning to reduce language-induced interference in semantic matching.

Abstract

Multilingual dense retrieval aims to handle queries and documents across different languages based on a unified retriever model. The challenge lies in enabling robust retrieval transfer to low-resource languages where annotated retrieval data is often scarce. Although previous studies transfer high-resource supervision to low-resource languages in multilingual semantic representation learning, the shared representation often entangles semantic and linguistic features, which may interfere with optimizing semantic relevance for retrieval. Different from existing methods that focus on learning language-agnostic semantic features under such entanglement, we propose a disentangled contrastive learning~(DCL) method for multilingual dense retrieval by separating multilingual representations into semantic and linguistic subspaces. Specifically, we design disentangled optimization objectives based on hierarchical semantic alignment and language debiasing contrastive learning. By aligning retrieval-relevant semantics across languages at both sentence and token levels while capturing language-specific variations in the linguistic subspace, these objectives reduce language-induced interference in semantic matching. We jointly optimize them with the retrieval objective to facilitate stable zero-shot transfer from English supervision to multilingual dense retrieval. Extensive experiments on mMARCO and MIRACL show that our method consistently outperforms several strong baselines, demonstrating its effectiveness and generalization ability.

View source

Similar papers

Book Open access Aug 2026

Semantic-Symbolic Knowledge Consensus for Multilingual Question Answering

This paper proposes SeSyCo, a Semantic-Symbolic Knowledge Consensus framework, which leverages the semantic space to diverge monolingual queries into broad multilingual evidence, and subsequently utilize the symbolic space to eliminate language discrepancies, converging the gathered information into a robust consensus...

Yu Zhang, Ran Song, Xiaofei Gao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence and achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.

Libiao Chen, Xiyang Liu, Yanheng Wei et al. · 1 citation
#natural language process... Preprint Oct 2026

Do Multilingual Encoders Produce Language-Consistent Semantic IDs?

Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we...

Abhinav Bohra, Anuj Bohra · 0 citations
Preprint Aug 2026

Cross-lingual Representation Learning via Centroid Intervention Fusion

Centroid Intervention Fusion is proposed, a projection fusion framework that consolidates multiple multilingual intervention projections into a single language-shared operator and outperforms the strongest prior pairwise intervention baseline by up to +3.3% across four model backbones.

Wei Sun, Marie-Francine Moens · 0 citations
Book Open access Aug 2026

LSAR: Sparse Lexical Representation Learning for Efficient and Interpretable Audio Retrieval

As Multimodal Large Language Models (MLLMs) expand the scope of retrieval-augmented generation, recommendation, and multimedia search, audio retrieval is expected to become a dependable retrieval component. Yet existing systems struggle to reconcile lexical precision, non-verbal acoustic evidence, and efficient, transp...

Haoyu Li, Yuzhe Bai, Li Niu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.