Skip to content

V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

Sep 2026 · 0 citations · 42 references
Computer Science Engineering

TL;DR

V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, is introduced and it is observed that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively.

Abstract

Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce V\={a}kQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. V\={a}kQA is publicly released.

View source

Similar papers

Conference Aug 2026

Vi-FusionQA: A Unified Benchmark and Baseline Study for Vietnamese Question Answering

Vietnamese question answering (QA) is increasingly important for customer support and digital service applications, yet existing Vietnamese QA resources remain limited in scale, domain coverage, and task format. In this paper, we introduce Vi-FusionQA, a public unified benchmark for cross-domain Vietnamese question ans...

Hong Son Nguyen, Duc Dat Pham, T. Huynh et al. · 0 citations
Open access 2026

Speaking ITM’s Language: Query Reformulation and Temporal-MMR for Long-Video QA

The resulting training-free pipeline, RECAST, consistently outperforms recent frame-selection baselines without modifying the ITM encoder, and preserves the per-frame matching cost of a standard single-query baseline.

S. Han, Thang Vu, Junyeong Kim · 0 citations
#artificial intelligence Preprint Sep 2026

SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning

Modern audio-language models are no longer judged only on what words they can transcribe, but on whether they can reason over what they hear: recovering meaning that lives in tone and prosody, telling dialects and regional languages apart, and resolving ambiguity that the written form leaves open. This capability is no...

Harshit Rajgarhia, Asif Shaik, R. Lokesh et al. · 0 citations
#large language models Open access Aug 2026

NormasTCU — A Brazilian Portuguese IR dataset and an evaluation of LLM-as-a-judge for relevance assessment

The results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

L. C. Fernandes, Marcus Vinicius Conceição de Castro, Leandro dos Santos Ribeiro et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.