A Lightweight Temporally Consistent Low-Light Enhancement Module for Visual-Inertial SLAM
Abstract
Low-light degradation weakens feature detection and inter-frame consistency in visual-inertial simultaneous localization and mapping. This paper proposes a lightweight temporally aware enhancement module that improves inter-frame illumination stability for downstream localization. At each time step, a causal three-dimensional depthwise separable convolutional predictor jointly uses the current and preceding frames to estimate the curve parameters of the current frame, which is then enhanced through iterative grayscale curve mapping. A structure-guided refinement module selectively enhances local edges while preserving the global illumination of the backbone output. The model is trained in two stages: the backbone uses zero-reference brightness and structure losses together with dark-region feature and temporal constraints, whereas the refinement module uses a pseudo-reference generated by fixed-parameter CLAHE solely for local detail guidance. Compared with Zero-DCE, the proposed method reduces global inter-frame brightness variation from 0.0049 to 0.0029 and increases ORB matches from 918.336 to 972.724, while processing 1280 by 720 images in 10.80 ms. Integrated into VINS-Fusion, it reduces the average absolute trajectory error by 19.6 percent relative to the unenhanced baseline on four sequences without loop closure.