Skip to content
Conference

Structured CT Imaging Artefact Assessment using Vision-Language Models

Jul 2026 · Signal Processing and Communications Applications Conference · pp. 1-4 · 0 citations · 17 references

Abstract

X-ray computed tomography (CT) is one of today’s most critical imaging modalities, with a wide range of applications spanning from medical diagnosis to industrial inspection. CT images can be severely affected by physically induced degradations such as low-dose noise and beam hardening, which compromise image quality and diagnostic accuracy. Existing AI-based approaches largely treat this problem as a pure classification task, and a system that explains the physical mechanisms of artefacts and provides actionable recommendations to the user has not been systematically addressed. In this study, we propose a vision-language model (VLM) based pipeline that detects CT artefacts, explains their physical mechanisms in natural language, and generates structured, actionable recommendations. LLaVA-1.5-7B and Qwen2-VL-7B models were fine-tuned using QLoRA on the 2DeteCT dataset; following fine-tuning, LLaVA-1.5-7B achieved 99.8% accuracy while Qwen2-VL-7B reached 86.9%. The results demonstrate the effectiveness of domain adaptation for structured artefact assessment.

View source