Deep Learning for Early Behavioral Screening of Autism Spectrum Disorder Applications Advances and Translational Challenges
Abstract
Deep learning is increasingly used to convert brief, observable behaviors into quantitative markers for autism spectrum disorder (ASD) screening. This narrative review synthesizes peer-reviewed work published from September 2021 to September 2026 on eye tracking, facial and body video, speech and language, response-to-name tasks, questionnaires, and multimodal digital phenotyping. Recent studies show a shift from static image classification toward standardized behavioral elicitation, temporal representation learning, and fusion of complementary signals. Prospective evidence is strongest for tablet-based computer vision and structured eye-tracking paradigms, where clinically relevant performance has been reported in young children; however, many studies remain limited by small or convenience samples, sample-level leakage, weak differential-diagnosis controls, and absent external validation. Apparent accuracy therefore depends as much on cohort construction and validation design as on architecture. Convolutional networks remain common for spatial features, while recurrent, temporal convolutional, and Transformer models increasingly represent behavior over time. Multimodal systems can improve coverage of both social-communication differences and restricted or repetitive behaviors, but they introduce missing-data, calibration, privacy, and interpretability challenges. The field is best positioned to deliver clinician-supported screening and triage rather than autonomous diagnosis. Progress requires subject-level evaluation, prospective multicenter cohorts, transparent reporting, uncertainty-aware outputs, fairness analysis, and alignment of model evidence with clinical criteria.