Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics, is introduced.
Abstract
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.
This work combines vocabulary surgery for tokenizer adaptation without full retraining, cross-lingual transfer of quality classifiers via translation, enabling filtering without target-language annotations, and generation of instruction data through translation, task conversion, and targeted synthesis.
Yurii Paniv, Bohdan Didenko, Mykola Haltiuk et al.· 3 citations
Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection...
An initial approach that first aligns heterogeneous table-cell representations into a shared space using Hirschfeld–Gebelein–Rényi maximal correlation (HGR) is proposed, and it is found that it generalizes competitively compared to models specifically designed for individual tasks.
This work introduces SteerWrite, a training-free framework designed for personalized co-writing that effectively adapts the base model to specialized domains without gradient updates, with specific designs tailored to small datasets.
Wenhao Mao, Chengbin Hou, Wei-Xiao Wang et al.· 0 citations
The findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through the multi-stage pipeline yields measurable gains in retrieval precision.
Savaş Yıldırım, Mucahit Cevik, Ayşe Başar· SN Computer Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.