Aug 2026· International Conference Computational Vision and Bio Inspired Computing· pp. 1212-1217· 0 citations· 21 references
Abstract
In this paper, a novel hybrid Swin Transformer and convolutional neural network (CNN) encoder-decoder, called StrokeDL-Net, is proposed for accurate multimodal brain MRI lesion segmentation and a temporal multi-modal fusion network, called ProgNet, is proposed for 90-day functional outcome prediction via the modified Rankin Scale (mRS). StrokeDL-Net overcomes the drawbacks of pure convolutional methods, especially for small and diffuse lesion detection, by combining local convolutional feature extraction with global shifted-window self-attention, augmented with an auxiliary penumbra estimation module. ProgNet fuses imaging biomarkers (segmentation), features from temporal evolution of MRI from acute and 24-hr follow-up scans, and structured clinical parameters using a cross-modal transformer attention mechanism. Extensive experimental evaluation on ISLES 2022, ATLAS v2.0, BraTS 2021, CQ500 and a proprietary clinical cohort (ISCHEMIA) achieves the Dice Similarity Coefficient (DSC) of 89.4%, sensitivity of 0.93, HD95 of 4.3 mm and mRS prediction accuracy of 87.3%, significantly outperforming all the baseline methods. The ablation studies verify the role of each component of the architecture. This framework constitutes a readily deployable clinical decision support system for acute stroke management. Beyond raw performance metrics, we validated StrokeDL-Net and ProgNet for clinical interpretability and generalizability across heterogeneous scanner vendors, field strengths and acquisition protocols, ensuring robustness to real-world data variability. Gradient-weighted class activation mapping confirmed consistent localization of model attention to clinically relevant infarct cores and penumbral tissue. Inference time was less than 12 seconds per case on average on standard clinical hardware, highlighting the practical feasibility of the framework to be integrated into the emergency stroke workflow.
This research systematically benchmarks five CNN architectures (VGG19, DenseNet201, ResNet50, Inception-v3, and MobileNet) on balanced and naturally imbalanced MRI datasets, suggesting that VGG19 is particularly good at discriminative performance.
Tegar Anugrah Firdaus, B. Rais, Marcelinus Jonathan Salim et al.· 0 citations
A novel hybrid convolutional neural networks and transformer architecture, HybCT-Net, augmented with a multi-level attention module and a regional explainability pipeline for brain tumor detection and classification is proposed, demonstrating superior performance than contemporary CNN, transformer and hybrid baselines.
Phanideep Karnati, Sukanya Roy, Dundi Urlamma et al.· International journal of com...· 0 citations
The results of the study reaffirm that the LES-based multimodal framework can provide an accurate, interpretable & computationally efficient diagnosis of early CVD & clinical decision support for clinical decision-making.
Indrapalli Swapna, Sasidhar Kothuru· International journal of com...· 0 citations
The proposed model such as MM-EffiFormer demonstrated the significant classification performance by achieving Accuracy of 0.990, Sensitivity of 0.987, Specificity of 0.995, Dice Similarity Coefficient (DSC) of 0.991, outperforms the existing model FCM-SVM.
Lovenish Sharma, S. Nanda· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.