Skip to content

Efficient Mixture-of-Experts for Video-Based Driver State and Physiological Multi-Task Estimation in Conditional Autonomous Driving

Oct 2026 · IEEE transactions on intelligent transportation systems (Print) · Vol 27, pp. 12096-12111 · 0 citations · 99 references

Abstract

Road safety remains a critical challenge worldwide, with approximately 1.19 million fatalities annually attributed to traffic accidents, often due to human error. As we advance towards higher levels of vehicle automation, challenges still exist, as driving with automation can cognitively over-demand drivers if engaging in non-driving-related tasks (NDRTs), or lead to drowsiness if driving is the sole task. This calls for the urgent need for an effective Driver Monitoring System (DMS) that can evaluate cognitive load and drowsiness in SAE Level-2/3 autonomous driving contexts. In this study, we propose a novel multi-task DMS, termed VDMoE, which leverages RGB video input to monitor driver states non-invasively. By utilizing key facial features to minimize computational load and integrating remote Photoplethysmography (rPPG) for physiological insights, our approach enhances detection accuracy while maintaining efficiency. Additionally, we optimize the Mixture-of-Experts (MoE) framework to accommodate multi-modal inputs and improve performance across different tasks. A novel prior-driven regularization method is introduced to align model outputs with human-factors priors, thus mitigating overfitting risks. We validated our method by creating a new dataset (MCDD), which comprises RGB video and physiological indicators from 42 participants, as well as two public datasets. Our findings demonstrate the effectiveness of VDMoE in monitoring driver states, contributing to safer autonomous driving systems. Code is in https://github.com/WJULYW/VDMoE

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.