Exploratory evaluation of large language models in patient-oriented orthopedic MRI report interpretation: a two-stage study
Abstract
Magnetic resonance imaging (MRI) reports for orthopedic conditions are filled with highly specialized terminology, leading to information asymmetry between clinicians and patients. Deficiencies persist in routine patient education regarding orthopedic imaging results, and a well-documented mismatch exists between physicians’ professional explanations and patients’ actual understanding of medical information. This study aimed to explore the feasibility of utilizing large language models (LLMs) to translate specialized orthopedic MRI reports into plain-language educational summaries, and to evaluate their impact on patients’ perceived comprehension and communication motivation. A prospective two-stage study design was adopted. In Stage I, two senior orthopedic surgeons evaluated 24 reports generated by six LLMs, assessing content accuracy via the Mika scale, reliability using the DISCERN tool, and objective readability. In Stage II, a total of 60 patients were randomly divided into three groups: conventional control group, Baichuan M3 group and Gemini 3.1 Pro group. We further compared changes in patients’ subjective comprehension and behavioral motivation after receiving AI-assisted interpretation. Stage I results showed that Baichuan M3 and Gemini 3.1 Pro delivered the best overall performance. Their accuracy (median Mika score = 1.0) and reliability were significantly superior to other models ( P < 0.001). Stage II demonstrated that baseline characteristics were comparable across the three groups ( P > 0.05). Following the AI intervention, both the Baichuan M3 and Gemini 3.1 Pro groups exhibited significant within-subject improvements compared to their respective baselines (all P < 0.001). No statistically significant difference was observed in the overall utility between the two AI models ( P > 0.05). Among all assessed domains, the dimension of action guidance and motivation demonstrated the most pronounced enhancement, with mean scores increasing by 2.5–2.8 points. LLMs show promise as auxiliary communication aids for patient-oriented orthopedic education. AI-assisted plain-language transformation is associated with higher patient-rated clarity and strengthened subjective motivation to engage in clinical dialogue. Given that imaging reports alone cannot replace comprehensive clinical reasoning, such tools should be deployed under strict safety guardrails to assist—not replace—physician-patient communication.