Aug 2026· SN Computer Science· Vol 7· 0 citations· 33 references
TL;DR
The findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through the multi-stage pipeline yields measurable gains in retrieval precision.
Abstract
This study examines the effectiveness of enhanced multilingual embedding models in improving retrieval performance for Turkish medical text data. We consider two specific medical applications in Turkish language: the TUS examination, a standardized medical assessment featuring exam questions, and Clinical QA, which involves authentic patient-physician interactions. By implementing multi-stage fine-tuning protocols on domain-specialized models, we provide detailed performance assessment and explore cross-domain transfer capabilities of the trained models. Our findings indicate that domain-specialized models improve in-domain retrieval relative to generic models, and that systematic optimization through our multi-stage pipeline yields measurable gains in retrieval precision. For instance, domain-specific fine-tuning improves TUS retrieval performance from 0.69 to 0.79 in P@1 and from 0.77 to 0.85 in MRR, while Clinical QA fine-tuning with hard-negative sampling improves P@1 from 0.33 to 0.39 and MRR from 0.41 to 0.48 relative to the vanilla encoder. In addition, reranking improves P@1 from 0.788 to 0.823 in our evaluated setting, corresponding to a 4.5% relative improvement. Furthermore, we find that, in multilingual model training, domain-specific knowledge acquired in the healthcare context of one language effectively transfers and enhances performance across other languages.
The main implication of this research is the validation of a practical architecture for improving IR systems, offering a viable alternative for domain-specific contexts such as Sequran.
Ray Ramadita, Wisnu Uriawan, W. Zulfikar· 0 citations
The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.
KoVRE: Korean Visual Document Retrieval Embedding is introduced, a single-vector retriever for Korean visual documents, alongside a comprehensive training recipe, demonstrating that targeted bilingual supervision and carefully designed training strategies can produce a highly effective Korean VDR model across diverse d...
This paper investigates the effectiveness of specializing ultra-compact language models for clinical embedding generation, utilizing the EmbeddingGemma 300M as the base model, and demonstrates the feasibility of performing fine-tuning on entry-level hardware.
The technical feasibility of text-based distractor selection rate prediction is established and the performance landscape across model categories is characterized, with potential application scenarios requiring future validation.
Zhe-Han Jiang, Tian-Peng Zheng, Jia-Yi Liu et al.· npj Digital Medicine· 0 citations
Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluati...
Nhan Phuoc Thanh Tran, P. Huynh, Trương Quốc Tuấn Trương et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.