This survey delivers a comprehensive and critical synthesis of the emerging role of GenAI across the autonomous driving stack, delving into the frontier applications of GenAI in image, LiDAR, trajectory, occupancy, and video generation, as well as LLM-guided reasoning and decision-making.
Abstract
Generative Artificial Intelligence (GenAI) constitutes a transformative technological wave that reconfigures industries through its unparalleled capabilities for content creation, reasoning, planning, and multimodal understanding. This revolutionary force offers the most promising path yet toward solving one of engineering’s grandest challenges: achieving reliable, fully autonomous driving, particularly the pursuit of Level 5 autonomy. This survey delivers a comprehensive and critical synthesis of the emerging role of GenAI across the autonomous driving stack. We delve into the frontier applications of GenAI in image, LiDAR, trajectory, occupancy, and video generation, as well as LLM-guided reasoning and decision-making. We categorize practical applications, such as end-to-end driving strategies and closed-loop simulations. We identify key obstacles and possibilities such as comprehensive generalization across rare cases, evaluation, safety, and onboard deployment. By unifying these threads, the survey provides a forward-looking reference for researchers, engineers, and policymakers navigating the convergence of generative AI and advanced autonomous mobility. An actively maintained repository of cited works is available at https://github.com/taco-group/GenAI4AD.
Artificial Intelligence (AI) has become the core enabling technology for Autonomous Driving (AD) systems, significantly improving perception, planning and control performance. However, the extensive adoption of deep learning and End-to-End (E2E) architectures has also introduced severe black-box problems, resulting in limited transparency and weak interpretability in safety-critical driving scenarios. The pursuit of transparent and accountable decision-making in autonomous driving has positioned Explainable Artificial Intelligence (XAI) as a pivotal area of inquiry in recent years. This paper systematically reviews the major black-box challenges in AI-based autonomous driving, including perception-level representation ambiguity, multimodal fusion interpretability issues, decision-making and planning unexplainability, and E2E driving model opacity. Furthermore, representative XAI techniques are analyzed from both technical and functional perspectives. The limitations and future research directions of XAI for AD are discussed. Through a systematic examination of existing studies, this review aspires to lay a solid foundation for the advancement of safer and more dependable AD systems.
Zhuo Song, Yao-Zong Zhang, Yu-Xiang Dong et al.· International Conference on...· 0 citations
It is argued that progress in vision generative AI has been driven by output quality, with hardware evolving reactively to accommodate growing model demands, making generative AI deployment sustainable and accessible across a much broader range of platforms.
Eleni Tselepi, Cristian Sestito, Shady O. Agwa et al.· 0 citations
The integration of artificial intelligence (AI) with mobile robotics is transforming automation technologies, enabling advancements across industries such as manufacturing, healthcare, and hazardous material handling. As AI-driven mobile robots become increasingly prevalent, their ability to enhance operational efficiency and safety underscores the need for a comprehensive understanding of their underlying methodologies. This study critically examines the integration of nature-inspired optimization algorithms and deep learning techniques in autonomous mobile robotics, addressing key research gaps related to adaptability, real-time decision-making, and energy-efficient operations. A significant contribution of this work is the exploration of the synergistic application of bio-inspired heuristics and AI-driven perception systems, which substantially improve environmental awareness and autonomous navigation. Furthermore, this review highlights existing limitations in computational complexity and ethical considerations within current frameworks, proposing a structured approach to mitigate these challenges. By systematically analyzing recent advancements and outlining future research directions, this study provides valuable insights that contribute to the development of more robust, intelligent, and ethically responsible autonomous robotic systems.
Ahmed N. Abdalla, J. K. S. Paw, Y. C. Tak et al.· Terra Joule Journal· 2 citations
A latent memory pool is constructed that stores failure cases along with their structure scene representations and expert trajectory labels, and a dedicated Retrieve Model that decouples static road structure and dynamic agent interactions to enable structurally grounded retrieval is designed.
Zebin Xing, Yupeng Zheng, Qiangyu Chen et al.· 0 citations
Although autonomous, large language model-driven systems show immense potential for orchestrating complex scientific experiments, their efficacy is constrained by two fundamental bottlenecks: context dilution, where strategic reasoning degrades as experimental history accumulates, and inter-campaign amnesia, which forces systems into computationally expensive tabula rasa explorations upon encountering novel domains. To overcome these limitations, we introduce the Multi-Objective State-Action Network (MOSAN), an autonomous cognitive information fusion framework designed for robust cross-campaign generalization and sensor shift adaptation. MOSAN achieves effective information fusion by coupling a self-evolving semantic memory with an Organic Strategic Graph Memory (OSGM), strictly orchestrated through a novel Single Ledger Architecture that isolates cognitive phases and prevents context saturation. A core mechanism of the OSGM is Cross-Domain Strategic Seeding (‘Ghost Node’ injection), a transient cold-start fusion strategy. Upon encountering novel datasets, the OSGM temporarily injects topological priors from similar past domains; these ephemeral nodes guide initial architectural deductions and are subsequently purged, allowing the system to disseminate structural knowledge across disciplines while decreasing the risk of cross-contamination. The fusion framework was validated on the challenging task of bacterial classification via Surface-Enhanced Raman Spectroscopy subject to sensor aging shifts, autonomously discovering highly accurate multi-topological architectures. Crucially, when subjected to an out-of-distribution 1D sequential modality (the ECG5000 cardiovascular dataset), the agent performed zero-shot architectural meta-learning, bypassing brute-force search to rapidly achieve 0.9930 accuracy. Finally, the framework was successfully deployed using an open-weight model (Mistral 24B) at zero API inference cost. By continuously fusing multi-epoch empirical evidence, bridging representational topologies, and actively overcoming sensor drift, MOSAN establishes a scalable and rigorous paradigm for autonomous biophysical discovery.
S. Domenico, B. Guilcapi, A. Milano et al.· Machine Learning: Science an...· 0 citations