Skip to content
#small language model Open access

From zero-shot to fine-tuning: optimize large language models for error detection of ultrasound reports

Sep 2026 · Insights into Imaging · Vol 17 · 0 citations · 29 references
Medicine

TL;DR

It is demonstrated that task-specific fine-tuning enables open-source LLMs to rival proprietary LLMs and expert radiologists in Chinese ultrasound report error detection, offering a locally deployable and privacy-preserving, and regulation-compliant solution for enhancing reporting consistency and reducing diagnostic errors in high-volume ultrasound examination procedures.

Abstract

High workload and inconsistent quality of ultrasound report writing may lead to diagnostic errors. This study aims to ascertain whether fine-tuned open-source large language models (LLMs) can achieve promising performance for automated quality control of Chinese ultrasound reports, when compared to proprietary LLMs. This retrospective, multi-center study included a multi-subspecialty dataset of 1800 Chinese ultrasound reports, comprising 1500 quality-controlled reports injected artificially with six predefined error types and 300 reports with naturally occurring errors. Nine proprietary LLMs (under zero-shot and few-shot paradigms) and seven open-source LLMs (under fine-tuning) were evaluated, with performance compared against that of radiologists of varying seniority. Performance was measured by detection accuracy, Macro-F1 score, precision, recall, and mean absolute error across six categories. Fine-tuned open-source LLMs, notably Qwen3-14B, achieved a detection accuracy of 0.931 and a Macro-F1 of 0.739, approaching the performance of senior radiologists. Some fine-tuned open-source LLMs maintained performance despite smaller parameter sizes and outperformed most proprietary LLMs with vastly larger parameter counts. The fine-tuned Qwen3-14B demonstrated superior recognition capability for semantic errors such as redundancy, spelling, orientation, and unit or value errors. This study demonstrates that task-specific fine-tuning enables open-source LLMs to rival proprietary LLMs and expert radiologists in Chinese ultrasound report error detection, offering a locally deployable and privacy-compliant alternative for AI-assisted clinical quality control workflows. Question Can task-specific fine-tuning improve error detection by open-source LLMs in Chinese ultrasound reports and provide an effective approach to automated report quality control? Findings Task-specific fine-tuning enabled open-source LLMs to achieve Macro-F1 scores up to 0.739, approaching that of experienced radiologists (0.764) and outperforming most proprietary LLMs. Critical relevance statement Task-specific fine-tuning allows open-source LLMs to become feasible assistants in ultrasound report quality control workflows, offering a locally deployable, privacy-preserving, and regulation-compliant solution for enhancing reporting consistency and reducing diagnostic errors in high-volume ultrasound examination procedures. Question Can task-specific fine-tuning improve error detection by open-source LLMs in Chinese ultrasound reports and provide an effective approach to automated report quality control? Findings Task-specific fine-tuning enabled open-source LLMs to achieve Macro-F1 scores up to 0.739, approaching that of experienced radiologists (0.764) and outperforming most proprietary LLMs. Critical relevance statement Task-specific fine-tuning allows open-source LLMs to become feasible assistants in ultrasound report quality control workflows, offering a locally deployable, privacy-preserving, and regulation-compliant solution for enhancing reporting consistency and reducing diagnostic errors in high-volume ultrasound examination procedures.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

It is suggested that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.

Hermione Warr, Harry Anthony, Lilli J. Freischem et al. · 0 citations
#machine learning Preprint Sep 2026

Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports

Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight m...

Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi et al. · 0 citations
Aug 2026

Improved Readability and Translational Instability in LLM-Generated Radiology Reports.

Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission, so optimized structured prompts can substantially reduce the variability of model outputs and improve the accuracy of medical text translation.

Yun Mao, Chunyan Wang, Wei Wang et al. · 0 citations
Open access Aug 2026

Vision-Language Model as a ‘Zero-Shot’ Assistant for Evaluating Condylar Osseous Changes in Cone-beam Computed Tomography

VLMs, particularly Gemini-3, exhibit zero-shot capabilities in identifying condylar changes and generating high-quality diagnostic reports, and demonstrate potential as preliminary, training-free screening aids in oral maxillofacial radiology.

Ke Chen, Andrew Zhang, Xian-Ju Xie et al. · 0 citations
Review Open access Aug 2026

Error Detection and Correction in Chinese Radiology Reports Using Large Language Models: Real-World Clinical Validation Study

Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.

Jia-Feng Zhou, Yuxin Wei, Qian Cai et al. · 0 citations
Open access Sep 2026

An Automated, Contamination-Controlled VQA Benchmark for Evaluating Vision-Language Models on 3D Oncology Imaging

Abstract Vision-language models (VLMs) are increasingly applied to medical imaging, yet public benchmarks may reward memorization over perception: their images and questions can enter pretraining corpora, and many items remain answerable from question text alone. We present an automated, agent-driven pipeline that buil...

Bo Liu, Han Gu, Xiang-Rui Li et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.