A Staged Workflow for Driver Glance and Hand-State Annotation
Abstract
Naturalistic driving video captures how drivers supervise under real road conditions, but converting long, multi-camera in-cabin footage into behavioral features still relies on manual coding, which is costly and hard to scale. We present a human-in-the-loop workflow that converts raw in-cabin video into two reviewable supervisory features: area-of-interest (AOI) gaze labels and hand-on-wheel state. Instead of a single end-to-end model, the workflow keeps region selection, gaze labeling, participant-specific calibration, temporal stabilization, and window-level metric extraction as distinct, observable stages, so an error can be localized. Specifically, the gaze module pairs Sample and Computation Redistribution for Efficient Face Detection (SCRFD) face detection with a YOLOv8-cls classifier, while the hand module uses GroundingDINO with a temporal state stabilization strategy. The workflow was validated in a naturalistic driving study with millions of video frames from 15 drivers. The source code of the workflow has been released at https://github.com/zdu881/autodri.