Efficient Monocular Depth Estimation on Embedded Systems with Neural Cellular Automata
Abstract
Real-time monocular depth estimation is an essential task for autonomous drone navigation, yet existing models require hundreds of billions of floating-point operations per in-ference, rendering them impractical for deployment on resource-constrained embedded systems. On the MidAir aerial imagery dataset, a mean absolute error of 5.05 m is obtained with mea-sured CPU inference latency of 3.66 ms and 40,832 parameters—2.9–6.1× faster and 5.4–16.3× fewer parameters than com-parable lightweight convolutional neural network (CNN) base-lines (FastDepth, MiniDepth, RT-MonoDepth-S, MiDaS-Lite). A depth-augmented semantic segmentation variant performs simultaneous depth estimation and semantic segmentation in a single forward pass with approximately 59,000 parameters, enabling holistic computer vision for autonomous flight. Diverse image and video processing tasks are supported by the same design, providing a practical foundation for efficient real-time computer vision in edge computing applications.