Skip to content
Preprint

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Aug 2026 · 0 citations · 27 references
Computer Science

TL;DR

Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics, is introduced.

Abstract

Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.

View source

Similar papers

Data-Efficient Adaptation of Multilingual LLMs to Ukrainian

This work combines vocabulary surgery for tokenizer adaptation without full retraining, cross-lingual transfer of quality classifiers via translation, enabling filtering without target-language annotations, and generation of instruction data through translation, task conversion, and targeted synthesis.

Yurii Paniv, Bohdan Didenko, Mykola Haltiuk et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Improving Cross-Lingual Token Representations by Adding a Pinch of SALT

Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection...

Guillem Ramírez · 0 citations

Can a single tabular embedding model service different tasks?

An initial approach that first aligns heterogeneous table-cell representations into a shared space using Hirschfeld–Gebelein–Rényi maximal correlation (HGR) is proposed, and it is found that it generalizes competitively compared to models specifically designed for individual tasks.

Unknown authors · 0 citations
Preprint Aug 2026

Training-Free Token-Level Steering for LLM Personalized Co-Writing

This work introduces SteerWrite, a training-free framework designed for personalized co-writing that effectively adapts the base model to specialized domains without gradient updates, with specific designs tailored to small datasets.

Wenhao Mao, Chengbin Hou, Wei-Xiao Wang et al. · 0 citations
Open access Aug 2026

Optimizing Multilingual Embedding Models for Retrieval and Reranking in RAG Pipelines: Enhancing Semantic Search in Turkish Medical Datasets

The findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through the multi-stage pipeline yields measurable gains in retrieval precision.

Savaş Yıldırım, Mucahit Cevik, Ayşe Başar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.