Configuring AI Glasses as Access Technology: Investigating How Blind and Sighted People Respond to Asynchronous Trouble in Image Description
Abstract
Wearables powered by computer vision and large language models (LLMs) such as AI glasses are increasingly marketed as assistive technology for blind and low vision people. However, HCI scholarship attending to AI-powered image descriptions draws attention to erroneous outputs, privacy issues, and the need to design automated systems for image description with blind and low vision experts. We analyze a corpus of video ethnographic material documenting how blind people inspect different environments with the Meta AI glasses alongside sighted people. Through transcriptions of video excerpts, we show how blind and sighted participants assist the technology in generating synthetic image descriptions. Building on the analysis, we develop the concept of asynchronous trouble to describe a spatiotemporal discrepancy where users’ actions and the glasses’ output do not share a common reference for description. Examining the consequences of asynchronous trouble, we offer empirical and methodological contributions to HCI research on automated image descriptions.