Skip to content

Author

Xiaolong Wu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Evaluation of a large language model for clinician-facing preoperative cost-communication preparation in total knee arthroplasty

Background and aims Costs associated with total knee arthroplasty (TKA) may affect treatment preparation, expectation management, and postoperative care planning. Previous large language model (LLM) studies have focused mainly on medical question answering, patient education, and clinical decision support, whereas their performance in clinician-facing preoperative cost-communication preparation remains unclear. This study evaluated an LLM using real-world clinical records from two hospitals in an expert-referenced offline evaluation. Methods Preoperative medical records of 80 patients who underwent primary unilateral TKA at two hospitals in China from January to May 2026 were included. A structured expert-panel process was used to develop a preoperative cost-communication framework comprising 4 dimensions and 17 clinical cues and to establish case-level minimum necessary communication items (CL-MNCIs) for each case. Task 1 assessed identification of the 17 cues. Task 2 assessed CL-MNCI coverage and classified all generated items according to case relevance, medical-record support, redundancy, and safety. Twenty-four cases were non-randomly selected by CL-MNCI count for three repeated-generation runs. Results Task 1 comprised 1,360 case–label classification units. Micro-precision, micro-recall, micro-F1, accuracy, MCC, and macro-F1 were 0.901, 0.888, 0.894, 0.948, 0.860, and 0.840, respectively. Experts established 569 CL-MNCIs, of which 483 were covered, yielding an overall coverage rate of 84.9%; complete coverage was achieved in 17 cases. The LLM generated 716 items, including 12 safety events across 9 cases, 483 items matching CL-MNCIs, 132 record-supported supplementary items, 50 redundant items, and 39 items with insufficient record support. Overall, 665 items (92.9%) were case-relevant and record-supported, although this proportion included redundant content. Pairwise Jaccard similarity for covered CL-MNCI sets ranged from 0.813 to 0.823. Of 180 CL-MNCIs, 126 (70.0%) were covered in all three runs, and case-level agreement in safety classification ranged from 87.5 to 95.8%. Conclusion The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained. Its use should be limited to clinician-reviewed communication preparation and should not replace direct patient cost disclosure or professional judgment.

Zebing Ma, Liping Xue, Gonghui Jian et al. · 0 citations