Skip to content
Open access

Optimizing Multilingual Embedding Models for Retrieval and Reranking in RAG Pipelines: Enhancing Semantic Search in Turkish Medical Datasets

Aug 2026 · SN Computer Science · Vol 7 · 0 citations · 33 references

TL;DR

The findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through the multi-stage pipeline yields measurable gains in retrieval precision.

Abstract

This study examines the effectiveness of enhanced multilingual embedding models in improving retrieval performance for Turkish medical text data. We consider two specific medical applications in Turkish language: the TUS examination, a standardized medical assessment featuring exam questions, and Clinical QA, which involves authentic patient-physician interactions. By implementing multi-stage fine-tuning protocols on domain-specialized models, we provide detailed performance assessment and explore cross-domain transfer capabilities of the trained models. Our findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through our multi-stage pipeline yields measurable gains in retrieval precision. For instance, domain-specific fine-tuning improves TUS retrieval performance from 0.69 to 0.79 in P@1 and from 0.77 to 0.85 in MRR, while Clinical QA fine-tuning with hard-negative sampling improves P@1 from 0.33 to 0.39 and MRR from 0.41 to 0.48 relative to the vanilla encoder. In addition, reranking improves P@1 from 0.788 to 0.823 in our evaluated setting, corresponding to a 4.5% relative improvement. Furthermore, we find that, in multilingual model training, domain-specific knowledge acquired in the healthcare context of one language effectively transfers and enhances performance across other languages.

Read PDF

Similar papers

Open access Aug 2026

Retrieval-augmented generation for medical question answering: a multi-metric performance evaluation

The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.

Yunus Kökver · 0 citations
Preprint Aug 2026

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval

KoVRE: Korean Visual Document Retrieval Embedding is introduced, a single-vector retriever for Korean visual documents, alongside a comprehensive training recipe, demonstrating that targeted bilingual supervision and carefully designed training strategies can produce a highly effective Korean VDR model across diverse d...

Yongbin Choi, Gyuho Shim, Youngjoon Jang · 0 citations
Open access Aug 2026

Transfer learning for text-based distractor selection rate prediction in medical multiple-choice questions: fine-tuning embedding models as a plausibility proxy

The technical feasibility of text-based distractor selection rate prediction is established and the performance landscape across model categories is characterized, with potential application scenarios requiring future validation.

Zhe-Han Jiang, Tian-Peng Zheng, Jia-Yi Liu et al. · 0 citations
Conference Aug 2026

Retrieval-Augmented Fine-Tuning with Reasoning Distillation for Vietnamese Medical Question Answering

Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluati...

Nhan Phuoc Thanh Tran, P. Huynh, Trương Quốc Tuấn Trương et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.