Skip to content
Preprint

Efficient Semantic Understanding from Digital Foveation

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

A lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation is introduced, suggesting that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively.

Abstract

Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.

View source

Similar papers

Open access Sep 2026

Exploring the Synergy Between Language Semantic Guidance and Visual Attention in Knowledge Distillation for Semantic Segmentation Under a Limited Field of View

Semantic segmentation for autonomous driving is challenged by limited field-of-view (FoV), where occlusions and restricted viewpoints reduce available visual information. Existing approaches often incur considerable computational overhead, which is unsuitable for autonomous driving systems. To build robust and lightwei...

Jin-Hyun Ryu, Seung-Woo Cha, Eunyeong Jeon · 0 citations
Preprint Sep 2026

HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding

Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-fidelity 3D maps. Most existing works focus on geometric reconstruction and lack semantic information for high-level spatial understanding and task planning. Also, as th...

Han-Wen Cao, Wen-Qiang Wu, Kuang-Ting Tu et al. · 0 citations
Preprint Sep 2026

CoordFormer: Give Me Any Coordinates and I Will Give You Labels

Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost and the difficulty of capturing fine-grained details. We propose CoordFormer, a novel coordinate-based architecture for semantic segmentation that predicts labels at arbitrary spatial locations through a Coordina...

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli et al. · 0 citations
Sep 2026

Real-time Robust Semantic Segmentation via Low-Resolution Ensemble.

Semantic segmentation algorithms constitute a critical component for vision-based perception in robotics and autonomous driving. Real-world deployment in uncertain environments requires such algorithms to be both low-latency and robust against common corruptions. However, the limited representation ability of real-time...

Yuanduo Hong, Hui-Hui Pan, Jue Wang · 0 citations
Open access Sep 2026

Intelligent Visual Prioritization for Retinal Prostheses via Context-Aware Object Ranking and Depth-Aware Phosphene Generation

Images from high-resolution cameras are mapped onto a sparse pattern of low spatial resolution and intensity in the retina, which limits visual perception in retinal prosthetic vision. When the entire scene is converted into phosphenes, it may allow unnecessary background information to be retained and may cause visual...

Xin-Wei Li, Irshad Khalil, Faisal Rahman et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.