Skip to content
Preprint

Dynamic Resolution Routing for Efficient Egocentric Grounding

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

This work proposes SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing and introduces a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall.

Abstract

Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable for selecting object-centric spatial evidence. To overcome this, we propose SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing. SmartRes first encodes a low-resolution view for global context and uses a lightweight router to activate high-resolution patches in object-centric regions and constructs an order-preserving visual sequence. To further enable robust routing under severe foreground-background imbalance, we introduce a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall. Experiments on Ego4D and EgoIntention show that SmartRes reduces visual tokens by up to 67% while retaining 86.4% of full-resolution performance, and achieves up to 1.66X faster inference than state-of-the-art token reduction methods with higher accuracy. Furthermore, strong performance on small object grounding indicates the effectiveness of SmartRes towards egocentric applications. Code will be publicly available.

View source

Similar papers

Conference Aug 2026

Region-Guided Search for Ultra-High-Resolution Small Object Detection

Current object detection models face significant challenges when applied to ultra-high-resolution images, as global downscaling frequently destroys crucial details for small objects, while EGC wastes computation on irrelevant regions and causes object truncation at tile boundaries. We propose a two-step reasoning archi...

Gia-Phuc Song-Dong, Anh-Kiet Tran-Nhu, Minh-Triet Tran et al. · 0 citations
Preprint Sep 2026

CoordFormer: Give Me Any Coordinates and I Will Give You Labels

Semantic segmentation on very-high-resolution images remains challenging due to the high computational cost and the difficulty of capturing fine-grained details. We propose CoordFormer, a novel coordinate-based architecture for semantic segmentation that predicts labels at arbitrary spatial locations through a Coordina...

Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli et al. · 0 citations
Preprint Sep 2026

Efficient Semantic Understanding from Digital Foveation

A lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation is introduced, suggesting that substantial semantic understanding can emerge from sparse observations when co...

Caterina Caccavella, Vittorio Fra, Andreas Ziegler et al. · 0 citations
Preprint Aug 2026

EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly scree...

Yifei Cao, Guolong Wang, Mingliang Hou et al. · 0 citations
Preprint Aug 2026

VidParse: Online Parsing of Egocentric Procedures Like a Pro

VidParse is presented, an online, training-free framework that treats activity understanding as a graph-constrained inference problem and achieves up to a 10x improvement in complex multi-step parsing accuracy over strong online baselines, all without requiring a single gradient update.

Anubhav Gupta, A. Kambhamettu, Vatsal Agarwal et al. · 0 citations
Conference Aug 2026

DG-SSR: Dual-Granularity Structured Scene Retrieval for Autonomous Driving

To address the inherent limitations of Vision-Language Models in long-tail object retrieval for autonomous driving, this paper proposes a Dual-Granularity Structured Scene Retrieval (DG-SSR) architecture. By decoupling text queries and visual features, we introduce a parameter-free mechanism that fuses local semantic s...

Nan Jiang, Tongxuan Xu, Yu-Jin Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.