Skip to content

Author

Huiting Liu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Mar 2026

Evaluating Large Language Model-Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study.

BACKGROUND Large language model (LLM)-powered virtual standardized patients (VSPs) enable scalable clinical skills practice, but the validity of AI-generated scores relative to faculty ratings remains unclear. OBJECTIVE To assess agreement between LLM-generated and faculty ratings of history-taking and communication performance, and to examine the influence of rater and case heterogeneity. METHODS In this cross-sectional study, 92 fourth-year medical students completed one of three 15-minute voice-based VSP cases (fever, diarrhea, cough). Ten blinded faculty raters scored performance (0-100 total; 0-50 domains). AI scores were generated by DeepSeek-V3 using a calibrated prompt. Agreement was evaluated using mixed-effects models, intraclass correlation coefficients (ICC[2,1]), Spearman correlations, mean absolute error (MAE), Bland-Altman analysis, and variance partition coefficients (VPC). RESULTS Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]). Rater variability accounted for 37% of residual variance in faculty total scores (VPC = 0.37). AI total scores were positively associated with faculty total scores (β = 0.37, 95% CI 0.26-0.48, P<.001; Spearman ρ = 0.50, 95% CI 0.34-0.65). Absolute agreement was moderate (ICC[2,1] = 0.51, 95% CI 0.34-0.65), with MAE of 3.11 points. Mixed-effects Bland-Altman analysis showed a non-significant mean bias (1.26 points) and 95% limits of agreement from -4.95 to 7.48 (width = 12.43 points), with proportional bias (β_mean = -0.55, P<.001). Agreement was stronger for information gathering (β = 0.46, ρ = 0.49, ICC = 0.54, VPC = 0.23) than for communication (β = 0.27, ρ = 0.28, ICC = 0.29, VPC = 0.52). A sensitivity analysis in the lowest quartile showed attenuated but consistent agreement (ICC = 0.38). CONCLUSIONS LLM-based scoring in a VSP showed moderate agreement with faculty ratings, performing better for information gathering than for communication. Due to rater and case heterogeneity, ceiling effects, and proportional bias, it is suitable for formative use and enhanced sampling in programmatic assessment, but not for independent high-stakes summative decisions. CLINICALTRIAL

Xiaoxing Gao, Xiaoming Huang, R. Hu et al. · 0 citations