Skip to content
Open access

Multi-head attention-driven multimodal feature integration network for autism spectrum disorder detection

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 35 references
Computer Science

TL;DR

Experimental evaluation on the proposed multi-head Attention-driven Multimodal Feature Integration Network framework demonstrates that the proposed MAtMFIN framework consistently outperforms state-of-the-art transfer learning and Deep Learning models, including a hybrid Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) model.

Abstract

The growing prevalence of Autism Spectrum Disorder (ASD) highlights the need for accurate and reliable intelligent screening systems for early behavioral assessment. However, ASD-related behaviors vary significantly across individuals, making diagnosis based on a single modality or subjective evaluation unreliable. Moreover, existing computational approaches often struggle to effectively model complex spatial and temporal dependencies in behavioral video data, leading to limited feature representation. To address these challenges, this study proposes a Multi-head Attention-driven Multimodal Feature Integration Network (MAtMFIN) that jointly analyzes Eye-Tracking (ET) scanpath data and behavioral video data to improve the robustness of ASD screening. The framework employs a Hierarchical Attention Refinement block (HARb) and a Spatial Enhancement Module (SEM) for effective feature refinement, along with a cross-modal cross-attention mechanism to capture complementary relationships between modalities. Experimental evaluation on the proposed multimodal ASD dataset demonstrates that the proposed MAtMFIN framework consistently outperforms state-of-the-art transfer learning and Deep Learning (DL) models, including a hybrid Convolutional Neural Network (CNN)-Long Short-Term Memory (LSTM) model, 2D-CNN with attention, LSTM-attention, Inflated 3D CNN, and Spatio-Temporal Graph Neural Networks (STGNN). The proposed method achieves a training accuracy of 94.6%, a testing accuracy of 93.2%, and an F1 score of 93.5%. These results indicate that effective cross-modal feature integration and attention-based modelling of spatio-temporal dependencies significantly enhance the ASD behavioral analysis, highlighting the potential of the proposed framework for intelligent healthcare applications.

Read PDF

Similar papers

Open access Aug 2026

A Multi-Modal Approach for Early Autism Recognition Using Deep Learning and Sensory Data

This project introduces a smart, hybrid system that combines advanced deep learning technology with proven treatment methods, aiming to close the gap between diagnosis and meaningful help for autism, by blending advanced computational analysis with trusted treatment practices.

S. Ahmed, Shaikh Faeik, Shaikh Israhil et al. · 0 citations
Jul 2026

Multibranch Attention and Fusion Network for Voice-Based Autism Detection Using Heterogeneous Audio Embeddings.

Voice-based analysis is attracting growing interest as a noninvasive means of identifying early markers of autism spectrum disorder (ASD). While pretrained audio models such as YAMNet and VGGish provide complementary views of children's speech, most existing studies rely on a single representation and do not explore how these embeddings may be combined in a structured manner. This work introduces the multibranch attention and fusion network (MBAFNet), an architecture designed to make fuller use of heterogeneous embeddings by processing each stream through its own convolutional encoder, extracting temporal cues at multiple scales, and modeling cross-representation interactions through a self-attention layer. A gating mechanism then regulates the relative contribution of each embedding before classification. Experiments conducted on the CASD-SC corpus under a subject-disjoint five-fold evaluation protocol show that MBAFNet achieves 94.17% accuracy, outperforming all evaluated baselines and previously reported state-of-the-art approaches. These findings indicate that carefully designed selective fusion can reveal complementary information contained in pretrained embeddings and supports the development of more robust speech-based ASD assessment frameworks, while broader validation remains necessary before screening-oriented use.

Mouad El Omari, Younes El Belghiti, Hanae Belmajdoub et al. · 0 citations
#machine learning Preprint Sep 2026

A multimodal large language model for evidence-based autism spectrum disorder screening

The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.

Jun Chen, Qi Zhao, Yun-Liang Jiang et al. · 0 citations
Open access Aug 2026

Improving Autism Diagnosis Across Ages Using Eye-Tracking and Temporal Transformer Models

Variation in gaze behavior due to age is currently a considerable challenge in building reliable eye-tracking systems for Autism Spectrum Disorder (ASD) diagnosis. However, existing strategies often focus on static gaze representation or dataset-based information, which can lead to limited generalization of findings depending on developmental groups and heterogeneous recording conditions. In this paper, we present a temporal transformer-based system for ASD classification using eye-tracking sequences. This allows you to model gaze behavior as a structured temporal process in the context of contextual attention, as well as employing entropy-based modeling for various distributions of variability over time and temporal consistency constraints to capture sequential gaze dynamics related to ASD behavioral patterns. The framework was evaluated using public eye-tracking corpus containing temporally ordered gaze recordings from ASD and TD participants across age groups. Five sequential experiments on baseline classification, class-balancing analysis, cross-age evaluation, ablation analysis, and cross-dataset transfer learning were performed to conduct experiment-based evaluations. Model performed 0.91 in in-domain Area Under the Receiver Operating Characteristic Curve (AUC) and 0.81 in F1-score on the primary eye-tracking dataset. In the cross-dataset assessment stage, the framework presented a relatively stable performance, with an AUC of 0.85 and an average F1-score of 0.74, irrespective of differences in participant distributions and recording conditions. Ablation analysis also revealed that entropy regularization and temporal consistency mechanisms played a significant role in model stability and classification performance. The ablation analysis provides additional insight into the contribution of the proposed framework components beyond the overall classification performance. Removing the entropy-based regularization reduced the model’s ability to represent variability in gaze allocation, whereas removing the temporal-consistency regularization resulted in less stable sequence representations during learning. These observations indicate that the proposed components complement the transformer-based sequence encoder by improving representation stability and preserving diagnostically relevant temporal information. Rather than acting as independent classifiers, the regularization mechanisms serve as supporting constraints that enhance the quality and robustness of the learned temporal representations. The results indicate that temporally structured gaze modeling is more robust, interpretable, and general in comparison to static gaze representations. In summary, the presented framework can represent a scalable and developmentally appropriate approach to gaze-based ASD classification and support the implementation of trusted neurodevelopmental screening systems.

Mohammed A. Alzain, M. Rokaya, D. Hemdan et al. · 0 citations
#explainable ai Open access Sep 2026

A hybrid deep learning framework for early autism screening

Abstract Objectives: Early diagnosis of autism spectrum disorder (ASD) remains a significant challenge due to the time-consuming and subjective nature of traditional diagnostic methods. This study proposes a reliability-oriented hybrid deep learning framework that provides a low-cost, scalable, non-invasive, and AI-assisted pre-screening tool for early ASD risk indication and referral support. Methods: The proposed system integrates two independent deep learning architectures: (1) a ResNet18 model optimized with 10-fold cross-validation using static periocular image data, and (2) a multi-CNN facial image classification ensemble combining ResNet50, EfficientNet-B0, and DenseNet121 architectures. The periocular pathway uses the fine-tuned ResNet18 fully connected softmax layer as its decision boundary. The facial pathway uses weighted probabilistic averaging, and the two subsystem outputs are subsequently combined through an OR-based reliability fusion rule. Additionally, explainable artificial intelligence (Grad-CAM) was employed to visualize decision-relevant regions. Results: The static periocular ResNet18 model achieved a sensitivity of 90%, while the multi-CNN facial classification ensemble reached a sensitivity of 87.1% and an AUC of 0.948. Under the conditional-independence assumption, OR-based reliability fusion yielded an analytically estimated system-level sensitivity of 0.9871, corresponding to a joint false-negative probability of approximately 1.29%. Conclusions: The developed hybrid model is positioned as an AI-assisted pre-screening and early-referral support tool, not as a replacement for clinical evaluation. By combining reliability-oriented parallel fusion with explainable AI, the system supports clinical transparency and offers a scalable pathway for early ASD risk indication.

H. Ünözkan, Hüseyin Ali Sarıkaya · 0 citations
Open access 2026

Detection of ASD Using DySTEP-GAT Architecture With EEG-Based Attention-Fused Brain-Net Graph Representation

Findings indicate that the learned brain network representations capture distinct and quantifiable abnormalities in ASD-related neural dynamics, establishing the proposed framework as a robust, interpretable, and measurement-aligned tool for objective ASD assessment.

Madhuparna Das, Poulomi Pal, M. Mahadevappa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.