Skip to content

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Jul 2026 · arXiv.org · Vol abs/2607.18704 · 0 citations · 28 references
Computer Science

TL;DR

A principal contribution of this work is the transparency-first framework, in which every reported metric is explicitly identified as measured, derived, or unavailable, thereby improving the traceability, interpretability, and reliability of speech analytics.

Abstract

Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searchable content through automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle generation. The system is built on a FastAPI backend with a real-time dashboard and adopts a three-layer architecture comprising (i) a transcription and diarization core based on Whisper-class automatic speech recognition and pyannote speaker diarization, (ii) an audio intelligence layer that extracts acoustic and linguistic features, including waveforms, spectrograms, pitch, speaking rate, silence, filler-word frequency, and sentiment, directly from the audio signal, and (iii) an integration layer that supports data export and downstream workflow integration. A principal contribution of this work is the transparency-first framework, in which every reported metric is explicitly identified as measured, derived, or unavailable, thereby improving the traceability, interpretability, and reliability of speech analytics. The paper presents the system architecture, benchmarking methodology, explainability and uncertainty framework, and key considerations for enterprise-scale deployment.

View source

Similar papers

Jul 2026

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

Body-worn camera (BWC) audio presents unique challenges including high ambient noise, variable recording conditions, and multiple overlapping speakers that make automated transcription and speaker labeling challenging. We propose a dual-path conversational intelligence framework that preprocesses raw BWC audio, separat...

C. Vishnoi, S. Khurana, A. Timmapur et al. · 0 citations
Preprint Aug 2026

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

SwanTale is proposed, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks, and adopts reward-conditioned quality control and Engram conditioning along with Unified MoE for multi-task and multi-audio-modality modeling.

Yu Zhang, Ruiqi Li, Changhao Pan et al. · 3 citations
Preprint Aug 2026

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.

Wen-Jun Huang, Q. Chu, Tiger Shao et al. · 0 citations
#natural language process... Preprint Aug 2026

Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

This work presents the first multilingual spoken hallucination benchmark comprising 12,013 news samples across English, Russian, and Kazakh with controlled hallucinations of three types and three severity levels, and assesses fine-tuned multilingual encoders and, in zero-shot in-context settings, multimodal decoder mod...

Meruyert Aristombayeva, Jason Samuel Lucas, Chaewan Chun et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.