Skip to content
Open access

Toward Standardized Interactivity: An AI-Enabled Adaptive Speaking Task for Computer-Delivered Large-Scale Assessment

Jul 2026 · Language Testing · 2 citations · 46 references

TL;DR

Two solutions are proposed: LLM-driven response evaluation using question-specific task completion rubrics for adaptive follow-up question selection and item response theory (IRT) scoring accounting for prompt variability to enhance adaptivity in an LLM-enhanced SDS-delivered interview task.

Abstract

Despite task interactivity eliciting different aspects of language use, more interactive, adaptive speaking tasks have been challenging to implement in computer-delivered large-scale high-stakes contexts. Recent advances in large language models (LLMs) offer a potential means to address this tension through spoken dialogue systems (SDSs) but fall short of full adaptivity. To enhance adaptivity in an LLM-enhanced SDS-delivered interview task, this article proposes two solutions: (1) LLM-driven response evaluation using question-specific task completion rubrics for adaptive follow-up question selection and (2) item response theory (IRT) scoring accounting for prompt variability. We examined the extent to which these solutions functioned to simulate adaptivity, from 5,909 test takers’ responses to an adaptive speaking task with an avatar and a non-adaptive monologic task. Automated task completion evaluations were comparable to expert ratings, though follow-up question evaluation showed low agreement both among human raters and between raters and the system. IRT scoring yielded higher test–retest reliability, and the adaptive task elicited more reciprocal language use than the monologic task. Unlike earlier SDSs that prioritized consistency over adaptivity, the real-time meaning-level evaluation with IRT modeling strikes a balance between standardization and adaptivity. These findings support the viability of adaptive speaking tasks in computer-based standardized assessments.

Read PDF

Similar papers

Book Open access Aug 2026

PLAI: A Pilot Study of Profile-Based Explanation for AI-Supported Learning

Large language models (LLMs) are increasingly used as on-demand conversational learning assistants, but they typically do not adapt explanations to a student’s background unless explicitly prompted. We present the Personalized Learning Assistant Interface (PLAI), a web-based prototype that generates explanations from l...

Furkan Ali Yurdakul, Yi-Man Wu, Maria Torres Vega et al. · 0 citations
Open access Sep 2026

Affect-aware conversational adaptation in mixed reality procedural tasks

A mixed-reality conversational system that integrates voice interaction, LLM-driven dialogue, and prosody-based affect-aware adaptation with a modular client–server architecture is presented, enabling emotion-related cues inferred from vocal prosody to be incorporated without interrupting conversational flow.

Andrea Antonio Cantone, Matteo Ercolino, M. Sebillo et al. · 0 citations
Preprint Aug 2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...

Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al. · 3 citations · ⚡1
#small language model Open access Sep 2026

Human versus machine

Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost, genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for...

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations
Open access Sep 2026

From direct scoring to cue integration: big five prediction from spoken responses to a picture-based self-projection task

Open-ended language tasks are increasingly used in personality assessment, but their psychometric value largely depends on two factors: the type of evidence a task elicits and how that evidence is represented before scoring. In this study, we examined five spoken responses to a picture-based self-projection...

Qi-Fan Yang, Jian Li · 0 citations
Review Open access Sep 2026

Personality-Adaptive Conversational AI for Emotional Support: A Simulation Study Integrating Big Five Detection with Zurich Model-Inspired Regulation

LLM-based conversational agents generate fluent responses but remain limited in adapting their supportive style to individual personality and emotional needs. We present a Detect–Regulate–Evaluate (D–R–E) architecture that performs turn-by-turn Big Five detection and applies Zurich Model-inspired behavioural regulation...

Duojie Jiahua, Samuel Devdas, Mirjam Stieger et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.