Skip to content
Preprint

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

Jul 2026 · 0 citations · 21 references
Engineering

TL;DR

This paper presents the tttAI system submitted to the TSA-ASR task of the SmartGlasses Challenge 2026, evaluated on both two- person dialogues and multi-party meetings, and proposes a cascaded architecture consisting of speaker diarization, overlap detection, target-speaker extraction, post-processing, and automatic speech recognition.

Abstract

This paper presents the tttAI system submitted to the TSA-ASR task of the SmartGlasses Challenge 2026, evaluated on both two-person dialogues (Track 1) and multi-party meetings (Track 2). The task requires time-stamped speaker-attributed speech recognition from smart-glasses recordings. This is particularly challenging due to long-form audio, multiple speakers, and frequent overlapping speech. We proposed a cascaded architecture consisting of speaker diarization, overlap detection, target-speaker extraction, post-processing, and automatic speech recognition. The diarization module extracts features via WavLM-Large, performs frame-wise speaker classification with a Conformer encoder, and then generates global speaker segments through embedding clustering. For overlapped regions, we apply a WeSep-based target-speaker extraction model with ECAPA-TDNN speaker embeddings. When the extraction is unreliable, a dominant-speaker fallback strategy is used. The final system uses FireRedASR2-AED with the first microphone channel. The submitted system has a total parameter count of approximately 1.53B. On Track 1, our system achieves a tcpCER of 7.10%. On Track 2, it achieves a tcpCER of 34.04% and ranks second on the leaderboard.

View source

Similar papers

Preprint Aug 2026

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging bec...

Dehui Gao, Zhixian Zhao, Zhennan Lin et al. · 1 citation
Review Jul 2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-sp...

Shuai Wang, Zihan Qian, Ke Zhang et al. · 1 citation
Preprint Sep 2026

VibeVoice-ASR-Streaming Technical Report

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency...

Yu-Jie Tu, Zhiliang Peng, Jianwei Yu et al. · 0 citations
Preprint Aug 2026

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

DiaScriber is proposed, an end-to-end multi-speaker diarization and transcription model built on a speech large language model that achieves superior performance over comparison methods across extensive multi-speaker scenario test sets and demonstrates outstanding generalization ability in unseen multi-speaker scenario...

Bing-Shen Mu, Xian Shi, Xiong Wang et al. · 0 citations
Conference Aug 2026

AI-powered meeting transcription and summarization system based on Jitsi Meet, Jigasi, and Vosk

An automated pipeline for transcription and summarization of video conferences held on the Jitsi Meet platform that requires no proprietary cloud API keys and is deployable on-premise via Docker Compose, making it suitable for organizations with strict data-privacy requirements.

G. Amirkhanova, L. Bektemir, Shyrailym Adilkyzy et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.