Multimodal Fusion in Physical AI: Hardware-Aware Strategies for Robust Perception on Embedded Autonomous Driving Platforms
Interaction with the physical world differentiates physical AI from other forms of AI. Autonomous driving exemplifies this; vehicles must perceive and respond to dynamic environments with human-like or better perception-reaction times. This survey addresses the fundamental challenge of deploying high-performance models for multimodal fusion in resource-constrained automotive environments. We organise state-of-the-art deep learning approaches into five paradigms—CNN-based, transformer-based, dense BEV-based, sparse-based, and hybrid—revealing trade-offs in accuracy, latency, and efficiency, as well as strengths and limitations in robustness under adverse operational design domains. The hardware-aware perspective is a differentiating contribution, presenting strategies for deployment on automotive platforms, reducing inference latency by up to 50% and improving robustness in adverse conditions by up to 20%. By synthesising sensor fusion, deep learning, compute platforms, and hardware-awareness, this work equips researchers and practitioners with actionable insights and strategies for perception systems, bridging theoretical advances and production-grade autonomous driving requirements.