Exploring ASD Identification Using Different Types of Eye-Tracking Features
Abstract
Eye tracking offers a non-invasive way of examining social visual attention in children and may support the objective identification of autism spectrum disorder (ASD). However, small-sample studies are often affected by individual differences, inconsistent feature selection, and limited integration of different gaze representations. This study explored a dynamic and static multiscale feature fusion framework for classifying children with pre-existing clinical ASD diagnoses and typically developing (TD) children using eye-tracking data collected during an emotion recognition task. After quality screening, 44 children aged 4–7 years were included (18 ASD and 26 TD). Static statistical features, dynamic sequences, early feature concatenation, and multiscale fusion were evaluated using participant-independent data partitioning, with static feature selection restricted to the training participants in each split. CNN–LSTM with multiscale fusion achieved an accuracy of 0.78 ± 0.08, ASD recall of 0.65 ± 0.22, F1-score of 0.71 ± 0.12, and AUC of 0.76 ± 0.16. A fully nested leave-one-participant-out evaluation provided additional internal validation of participant-level classification. Overall, the results indicate that different approaches to combining static and dynamic gaze representations may lead to different classification outcomes, while further validation in larger independent samples is needed.