Skip to content
Preprint

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

Aug 2026 · 1 citation · 66 references
Computer Science

TL;DR

This work presents CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement, and reports quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.

Abstract

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters is not task success but whether the owner would endorse how their proxy handled the call. We evaluate only the text-domain conversational decision layer; speech recognition, audio interaction, end-to-end latency, and handset execution are outside scope. We present CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement. Each is paired, where available, with a counter-metric and an uncertainty estimate; no benchmark-wide Q1-Q5 composite or leaderboard score is defined. Three guardedness diagnostics identify candidate cases for a toolless proxy that holds no credentials and calls no tools. Across three model families represented by paired 4-bit checkpoints (0.6-4B), the primary scoring snapshot gives the larger checkpoint higher point estimates on several service, recall, and plausibility measures, while triage discrimination follows a different ordering. Bare scam-side TPR rewards universal suspicion, and pairwise separation changes when legitimate-side false positives are included and across judge snapshots. Scripted degenerate agents expose further floors, including a hangup-and-echo policy with entity recall 1.000. We report quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.

View source

Similar papers

#artificial intelligence Review Sep 2026

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations...

Pritish Mishra, Ishaan Kumar, Akshat Mandoli et al. · 3 citations
#natural language process... Preprint Sep 2026

Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model

Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this read...

Si-Miao Ren, Kidus Zewde, Xing-Yu Shen et al. · 6 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Talk2Agent: Benchmarking Voice Interfaces for Text Agents

Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or ta...

T. Chiba, Guang-Zhi Sun, Zhe-Qi Yuan et al. · 0 citations
#natural language process... Preprint Sep 2026

Where a Model Sends Its Own Repeated Token

Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax...

Nicolás Vera Zúñiga · 0 citations
Conference Aug 2026

NEMO: AI Assisted Deskbot

NEMO is an intelligent desktop assistant that performs tasks via voice commands in natural language with an average response time of 1.2-1.8 seconds and a command accuracy of 92-95%. It features a wake-word detector, real-time speech-to-text transcription using Google's Speech API, and a dual-mode natural language unde...

Trupti Kudale, Aditya Mhaisdhune, Atharva Mulane et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval (VLME) benchmark, where every answer depends on paralinguistic metadata attached to conversational turns, which is otherwise unrecoverable from the words alone, is presented.

Ramit Pahwa, Parivesh Priye, Apoorva Beedu · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.