Environment-aware dynamic prompting for self-supervised monocular depth estimation
Abstract
Self-supervised monocular depth estimation (MDE) eliminates the reliance on expensive ground-truth depth annotations and has emerged as a powerful approach for a wide range of vision applications. However, current lightweight networks are hampered by two critical challenges: the limited representation capacity of static weights in complex driving environments, and the propagation of noise through conventional skip-connections. To address these issues, we propose ScenePrompt-Mono, a lightweight architecture built upon the Lite-Mono backbone and driven by environment-aware dynamic prompting. Specifically, it integrates the Dynamic Context Prompt Generator (DCPG) to adaptively modulate encoder features for scene-specific variations, together with the Prompt-Guided Filter (PGF) to purify cross-level representations and preserve crisp geometric boundaries. Experiments on the KITTI dataset demonstrate that ScenePrompt-Mono achieves a competitive accuracy-efficiency trade-off.