Aug 2026· Frontiers in Medicine· Vol 13· 0 citations· 40 references
Medicine
TL;DR
A novel two-stage deep learning architecture that integrates self-supervised representation learning with supervised classification for kidney CT image analysis using a publicly available kidney CT image dataset is introduced.
Abstract
Introduction Kidney-related disorders are one of the global health concerns that require timely detection to prevent severe health complications. The use of computed tomography (CT) images for accurate classification of kidney diseases is important. However, it is challenging to differentiate between classes due to the subtle visual differences. This study introduces a novel two-stage deep learning architecture that integrates self-supervised representation learning with supervised classification for kidney CT image analysis using a publicly available kidney CT image dataset. Methods In the first stage, the DINO framework with a Data-efficient Image Transformer (DeiT-Tiny) backbone is used to learn useful features from kidney CT images independent of labels. In the second stage, the pre-trained model is fine-tuned using labeled data to classify kidney abnormalities. To ensure model transparency and clinical trustworthiness, two explainable AI techniques are applied. Grad-CAM++ is used to highlight important regions contributing to predictions in kidney CT images. In addition, DINO’s inherent multi-head self-attention mechanism is analyzed across all attention heads to capture diverse attention patterns. Results and discussion Experimental findings indicate that the proposed framework achieves strong classification performance, with a test accuracy of 99.16%, AUC-ROC of 99.99%, F1 score of 98.97%, precision of 98.90%, and recall of 99.05%, while also providing clear interpretability for automated kidney disease classification. External validation on a CT dataset from Iraq has yielded 97.03% accuracy, supporting the generalizability of the proposed framework.
A hybrid architecture in which ResNet50 is employed for localized spatial feature extraction, while Vision Transformer enables global contextual learning to automatically classify kidney tumors into multiple classes is proposed.
Hivi Kamal, W. M. Abduallah· passer of basic and applied...· 0 citations
Due to its late identification and challenging diagnosis, lung cancer continues to be one of the top causes of death for cancer patients globally, positioning it as one of the most critical concerns. Timely identification of cancerous nodules is essential for enhancing the patient’s survival likelihood CT image analysi...
S. Jegadeesan, S. Matheswaran, R. Palanivelrajan· International Conference on...· 0 citations
However, early-stage lung cancer is one of the major causes of death from cancer due to its symptoms being absent or hard to detect by conventional clinical analysis. The late diagnosis of this disease causes the effectiveness of the treatment to become poor and low survival rates. Traditional interpretation of medical...
R. Dhamotharan, T. Manikumar· International Conference Com...· 0 citations
The results suggest that the hybrid CNN-Transformer model provides a strong level of diagnostic accuracy and meaningfully understood visual rationale so it can serve as an excellent decision support mechanism for hospitals and radiologists in their daily operations.
Prasanna Pabba, N. S. Chaitanya, M. Ravikanth et al.· Journal of Intelligent Decis...· 0 citations
One of the main causes of cancer-related fatalities globally is lung cancer, and increasing patient survival requires early diagnosis. Lung cancer screening frequently uses computed tomography (CT) imaging, although manual interpretation can be laborious and reliant on radiologist skill. A deep learning-based framework...
Belayet Hossen, Md Takbir Alam Manjar, Md Abu Shihab et al.· International Conference Com...· 0 citations