Skip to content

Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision

Sep 2026 · 0 citations · 12 references
Computer Science

TL;DR

This submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge is presented, which ranked first in the large-model division and second in the<=2B division, suggesting that visual grounding is more important than annotation volume for this task.

Abstract

We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the<=2B division. The task requires a wearable assistant to decide after each eight-second segment of egocentric video whether to intervene or remain silent. Our approach has two main components. First, we reformulate intervention timing as single-token classification. Rather than generating either $interrupt$or $silent$, the model predicts yes or no, and we derive the decision from the renormalised probabilities of these two tokens. This formulation improved macro-F1 by 0.249 and G-mean by 0.30 over free-form generation. Second, because labelled data were limited to the released validation set, we generated additional supervision using a tool-calling video agent that inspects each clip and assigns intervention timestamps. A narration-only alternative was four times larger and ten times cheaper, but transferred worse than supervision from an unrelated real corpus, suggesting that visual grounding is more important than annotation volume for this task.

View source

Similar papers

Preprint Aug 2026

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when...

Yu-Chien Tang, Yuhan Liu, An-Zi Yen · 0 citations
Preprint Sep 2026

EgoTSR++: Egocentric Spatiotemporal Reasoning for Task Progress Understanding

Vision-Language Models (VLMs) have advanced rapidly in static visual understanding, yet remain unreliable when judging how an egocentric task is progressing. Given a task instruction and two visual observations, a model should determine which state is closer to the goal by analyzing task-relevant object configurations...

Xiao-Da Yang, Can Wang, Yu-Xiang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Companion-style QA Assistance in Ego-Vision

AI companions are envisioned as always-on assistants that support users in daily life. With this regard, we introduce BuddyVQA, a benchmark for companion-style question answering (QA) on egocentric streaming video. BuddyVQA contains 21.6K questions linked to 6K highlight moments across 1,012 long, egocentric videos. It...

Hangyu Qin, Jun-Bin Xiao, Sheng Zhang et al. · 0 citations
Aug 2026

Anticipating Object Interactions Via Aggregation and Distillation of Spatio-Temporal Knowledge From Vision Language Models.

ST-KAD sets a new state of the art, demonstrating accurate what-when-where prediction of future interactions, and confirms that the prior-informed aggregation and teacher-student distillation generalize beyond anticipation to spatial localization, validating the generality of the design.

Yang Liu, Dejie Yang, Minghang Zheng et al. · 1 citation

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.