Aug 2026· Academic Radiology· 0 citations· 21 references
Medicine
TL;DR
Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission, so optimized structured prompts can substantially reduce the variability of model outputs and improve the accuracy of medical text translation.
Abstract
Background
Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application.
Objectives
To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability.
Materials And Methods
This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education.
Results
Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics.
Conclusion
Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.
LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.
A. Serapio, Timothy L. Chen, Brian Tangsombatvisit et al.· Radiology· 1 citation
Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.
Jia-Feng Zhou, Yuxin Wei, Qian Cai et al.· Journal of Medical Internet...· 0 citations
PURPOSE
Radiology and pathology reports are important in breast surgical oncology planning, but are often unstructured and variable, requiring time-intensive previsit review and preparation. Here we evaluate the accuracy, completeness, and safety of retrieval-augmented large language model (LLM)-generated structured pr...
T. Moo, Robert James, Solange Bayard et al.· JCO Clinical Cancer Informat...· 0 citations
LLM-based patient-facing pathology report interpretation shows potential to bridge specialist pathology language and patient communication, and evaluation should extend beyond readability to include fidelity to the original pathology report, patient understanding, safety, and usability.
Chen Wang, Jie Hao, Si-Jia Zhang et al.· International Journal of Med...· 0 citations
Background: Artificial intelligence (AI) tools are being used to translate radiology reports into plain language, but translation errors may compromise comprehension and safety. Objective: To develop and evaluate a rubric for assessing the quality and safety of AI-generated patient-friendly radiology reports. Methods:...
Bonnie A. Armstrong, Arogya Koirala, Hye Sun Na et al.· AJR. American journal of roe...· 0 citations
This review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.
Fabio Galbusera, Andrea Cina· European spine journal· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.