Skip to content
Preprint

An active-learning framework for real-time depth perception from monocular vision streams

Aug 2026 · 0 citations · 44 references
Computer Science

TL;DR

Experimental results demonstrate that adaptability is not solely determined by model size, but rather by how effectively parameter plasticity is regulated in dynamic environments.

Abstract

Biological visual systems can perceive depth from monocular vision flow, continuously integrating temporal visual cues while maintaining a balance between stability and plasticity in dynamic environments. In contrast, artificial perception models deployed on resource-constrained edge devices are typically trained in a static offline manner and remain frozen after deployment, often suffering severe performance degradation under domain shifts. While large-scale models may encode broad knowledge through massive parameter redundancy, lightweight networks face a static optimization dilemma: forcing compact models to learn universal geometric representations is computationally inefficient and often leads to performance saturation. To resolve this issue, an Online Active Learning (OAL) mechanism is introduced to endow compact neural networks with the capability to adapt continuously during operation. A closed-loop Predict-Evaluate-Correct learning paradigm is established to actively select high-confidence, information-rich signals from streaming visual input. Crucially, Elastic Weight Consolidation (EWC) is employed not merely to prevent catastrophic forgetting, but to enforce Selective Plasticity, preserving parameters that encode globally relevant structural knowledge while allowing local alignment to newly observed environments. Built upon a MobileNetV3-Small backbone, the proposed system achieves approximately a 75% reduction in computational cost while maintaining competitive depth estimation accuracy. Experimental results demonstrate that adaptability is not solely determined by model size, but rather by how effectively parameter plasticity is regulated in dynamic environments.

View source

Similar papers

Review Open access 2026

Fusion-Oriented Deep Learning-Enhanced Visual SLAM: A Review

This paper presents a systematic review of deep learning-enhanced VSLAM, with a particular focus on how learning models are incorporated into classical simultaneous localization and mapping (SLAM) pipelines and how they function within the overall system.

Xiruo Chen, Qi Ouyang, Sihong Meng et al. · 0 citations
Conference Aug 2026

Lightweight Gated Residual Late Fusion for Frozen Vision Priors in Vision-Only Diffusion Policies

Visuomotor diffusion policy has shown strong capability in robotic manipulation, yet its performance depends heavily on visual representation quality. Under vision-only settings, limited expressiveness of conventional encoders bottlenecks both policy performance and convergence. Large-scale self-supervised vision found...

Jun-Lai Li, Qian-You Zhao, Duidi Wu et al. · 0 citations
Preprint Aug 2026

Falcon Perception-HD: High Density Perception via Reinforcement Learning

This paper explores post-training reinforcement learning (RL), specifically GRPO, to directly align autoregressive perception models with their evaluation metrics, and designs an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control.

Sofian Chaybouti, Yasser Dahou, N. Huynh et al. · 0 citations
Aug 2026

Adaptive multimodal visual tracking via parameter-efficient vision transformers.

Multimodal visual object tracking (MVOT) is crucial for achieving robust performance in complex environments, including scenarios with occlusion, low illumination, or high-speed motion. However, current methods often fall short in two aspects: their reliance on static fusion strategies limits adaptability to varying sc...

Yi-Xin Xu, Wenkang Zhang, Tian-Yang Xu et al. · 0 citations
Preprint Sep 2026

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

CAP is proposed, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information.

Hong-Jin Chen, Zi-Jun Xu, Shi-Hao Ma et al. · 0 citations
Open access Aug 2026

SLEA: a stochastic saccadic lightweight efficient attention framework with MobileNetV2 for robust and explainable 2D medical image classification

This work introduces saccadic lightweight efficient attention (SLEA), a framework purpose-built for 2D medical image classification and inspired by selective visual processing in the human visual system, intended to reduce dependence on fixed image locations and recurring spatial artefacts.

Williams Ayivi, Xiao-Ling Zhang, A. Aligayev · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.