Skip to content
Conference

Real-Time Multilingual Speech-to-Text AR Captioning Glasses for the Deaf and Hard-of-Hearing

Aug 2026 · 2026 International Conference on Secure Information Systems and Technologies (ICSIST) · pp. 219-225 · 0 citations · 19 references

Abstract

For people who are deaf or hard-of-hearing (HoH), everyday conversations can be difficult to navigate, even with modern hearing aids. Augmented reality (AR) smart glasses offer a practical way to provide live captions directly in the user’s field of view. However, current models are often too expensive, rely exclusively on strong internet connections, or fail to handle multiple, code-mixed regional languages. In this research work, we present a lightweight, fully functional pair of AR smart glasses designed specifically to provide real-time, multilingual speech-to-text with speaker identification and live translation. The glasses capture audio using an I2S MEMS microphone and stream it in 30 ms chunks to a custom React Native smartphone application. From there, the audio is routed to a tri-node AI engine. For global languages, the app streams to the Deepgram API. To specifically address the linguistic diversity of India, it utilizes the Sarvam AI WebSocket for ultra-low latency transcription and instant English translation of code-mixed Indian languages (e.g., Tamil, Hindi). If the network drops, the app seamlessly falls back to a locally stored 39-million-parameter Whisper AI model running entirely offline on the phone’s CPU. Our experiments show the cloud pipelines deliver translated captions to the OLED heads-up display (HUD) in roughly 420-450ms, while the offline model takes about 1850ms.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.