Skip to content
Open access

Transfer learning for text-based distractor selection rate prediction in medical multiple-choice questions: fine-tuning embedding models as a plausibility proxy

Aug 2026 · npj Digital Medicine · 0 citations

Abstract

Distractor selection rates in multiple-choice questions (MCQs) provide a behavioral proxy for distractor plausibility, yet current evaluation methods are reactive and resource-intensive. We investigated whether domain-specific fine-tuning of pre-trained embedding models can predict distractor selection rates from textual features before test administration. Using 6000 medical MCQs across eight clinical disciplines, we evaluated five general-purpose and five medical domain-specific embedding models under a unified 5-fold cross-validation protocol. Fine-tuning produced substantial improvements across both model categories: among medical models, SapBERT improved from r  = 0.403 to r  = 0.644 (+59.9%), while BGE-large improved from r = 0.467 to r  = 0.626 (+34.0%) within the general group. Compared with lexical baselines where TF-IDF with string overlap features achieved the best performance ( r  = 0.546), the proposed transfer learning with fine-tuned contextual models showed meaningful improvement. Meanwhile, compact models also performed competitively: MiniLM (22 M parameters) reached r  = 0.627 and MedEmbed-small (33 M) reached r  = 0.629. These results establish the technical feasibility of text-based distractor selection rate prediction and characterize the performance landscape across model categories. This article offers a methodological investigation of a plausibility proxy, with potential application scenarios requiring future validation.

Read PDF