Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Transfer learning for text-based distractor selection rate prediction in medical multiple-choice questions: fine-tuning embedding models as a plausibility proxy

Distractor selection rates in multiple-choice questions (MCQs) provide a behavioral proxy for distractor plausibility, yet current evaluation methods are reactive and resource-intensive. We investigated whether domain-specific fine-tuning of pre-trained embedding models can predict distractor selection rates from textual features before test administration. Using 6000 medical MCQs across eight clinical disciplines, we evaluated five general-purpose and five medical domain-specific embedding models under a unified 5-fold cross-validation protocol. Fine-tuning produced substantial improvements across both model categories: among medical models, SapBERT improved from r  = 0.403 to r  = 0.644 (+59.9%), while BGE-large improved from r = 0.467 to r  = 0.626 (+34.0%) within the general group. Compared with lexical baselines where TF-IDF with string overlap features achieved the best performance ( r  = 0.546), the proposed transfer learning with fine-tuned contextual models showed meaningful improvement. Meanwhile, compact models also performed competitively: MiniLM (22 M parameters) reached r  = 0.627 and MedEmbed-small (33 M) reached r  = 0.629. These results establish the technical feasibility of text-based distractor selection rate prediction and characterize the performance landscape across model categories. This article offers a methodological investigation of a plausibility proxy, with potential application scenarios requiring future validation.

Zhehan Jiang, Tianpeng Zheng, Jiayi Liu et al. · 0 citations