Skip to content
Review

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

Aug 2026 · 0 citations
Computer Science

TL;DR

This survey is the first to systematically study smart glasses through a unified framework, formalizing first-person data flow and constrained task utility, and introducing an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling.

Abstract

Smart glasses are evolving from capture and display accessories into first-person intelligence platforms that connect human perception, persistent context, and digital or physical action. Their on-body viewpoint aligns with the wearer's vision, audition, motion, and hand-object interaction, but must operate under tight energy, thermal, privacy, and feedback constraints. Despite rapid progress in augmented reality, egocentric vision, multimodal models, human-computer interaction, and embodied intelligence, the literature remains fragmented across devices, tasks, and benchmarks. \textit{The key challenge is not whether a model can recognize, answer, remember, or act in isolation, but whether a complete system can sustain a reliable, temporally valid, correctable, and governable perception-state-interaction-action loop.} This survey is \textit{the \textbf{first} to systematically study smart glasses through such a unified framework}. We formalize first-person data flow and constrained task utility, characterize devices along eight verifiable hardware capability axes, organize the literature around seven interdependent foundational capabilities, and introduce an L0-L5 framework spanning capture, reactive perception, contextual assistance, persistent state, governed action, and embodied coupling. Across nine application scenes, we connect tasks with datasets, systems, products, stakeholders, failure consequences, and evidence gaps. We further present a nine-dimensional deployment framework, a claim-conditioned evaluation protocol, and an evidence ladder from controlled measurement to longitudinal field validation and audit. Together, these elements make smart glasses more comparable, deployable, and reproducibly evaluated, while outlining a roadmap toward trustworthy first-person intelligence.

View source

Similar papers

Review Sep 2026

AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization

Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints....

Xu Yuan, Yi Wang, Zhuohang Jiang et al. · 0 citations
Preprint Sep 2026

Track, Articulate, Act: Generating Articulation from Casual Human Videos

Human videos contain rich causal evidence for robot manipulation: they reveal how hand motion induces object motion and produces task-relevant changes in object state. In this work, we study articulated objects such as doors, drawers, cabinets, laptops, ovens, and hinged containers that are ubiquitous in daily life and...

Jia-Ming Zhang, Homanga Bharadhwaj · 0 citations
Preprint Sep 2026

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedur...

Zhong-Bo Zhang, Jia-Yi Jin, Yi-Fan Wang et al. · 0 citations
Book Open access Oct 2026

Configuring AI Glasses as Access Technology: Investigating How Blind and Sighted People Respond to Asynchronous Trouble in Image Description

Wearables powered by computer vision and large language models (LLMs) such as AI glasses are increasingly marketed as assistive technology for blind and low vision people. However, HCI scholarship attending to AI-powered image descriptions draws attention to erroneous outputs, privacy issues, and the need to design aut...

B. Carreras, Damien Rudaz, Brian L. Due et al. · 0 citations
Preprint Aug 2026

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence, contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alp...

Fei Ma, Ze-Bang Cheng, Ming-Hui Li et al. · 0 citations
Preprint Aug 2026

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

This work presents Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis, and proposes OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding.

Liang Xu, Cheng-Qun Yang, Zili Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.