Aug 2026· International Conference on Multimedia Analysis and Pattern Recognition· pp. 19-24· 0 citations· 11 references
Abstract
Medical question answering (QA) plays a crucial role in clinical decision support, yet robust performance requires models to effectively distinguish relevant evidence from topically similar distractors within retrieved contexts. Existing Vietnamese medical QA benchmarks, however, focus exclusively on zero-shot evaluation, leaving this capability underexplored. In this work, we propose a retrieval-augmented fine-tuning (RAFT) pipeline for Vietnamese medical QA that explicitly teaches models to filter distractors and extract clinically relevant evidence. Our approach distills Chain-of-Thought (CoT) rationales from a large language model into a smaller model via Low-Rank Adaptation (LoRA), using oracle-and-distractor contexts constructed from VMHQA. To enable precise evaluation of reading comprehension independent of retrieval quality, we further introduce a fixed oracle-context protocol that isolates context utilization ability. Experimental results show that our fine-tuned LLaMA-3-8B achieves 93.64% accuracy and 93.56% macro-F1, outperforming the provided-context baseline by 18.76% and surpassing zeroshot LLaMA-3-70B. Evaluated under these optimal evidence conditions, our model demonstrates reasoning capabilities highly competitive with prior results on VMHQA, including fine-tuned GPT-4o (90.10%). Qualitative analysis further indicates that CoT distillation reduces context-induced hallucinations and improves robustness to distractors.
The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text g...
This submission to the MedReason 2026 challenge is described, covering multiple-choice (MCQ) and open-ended (OE) medical visual question answering (VQA) under fully offline, containerized inference, and it is found that MCQ retrieval must compare answer \emph{semantics} rather than answer labels.
Tristan Kirscher, N. Koser, Sören Pirk· 0 citations
This work proposes HybridRAG-BN, a retrieval-augmented framework for Bangla KBQA that integrates hybrid retrieval using BM25 and BGE-M3, answer generation using the GGUF version of Gemma-4-31B-Instruct, and a LoRA-fine-tuned Gemma-4-31B-Instruct model for answer verification and refinement.
This work studies whether a carefully domain-adapted retrieval-augmented generation pipeline closes the gap between compact and compact model quality in financial institutions under dense, frequently amended rulebooks.
Tobias Deußer, Abhishek Pillai, A. Bariviera et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.