Skip to content

CallScreenBench: Benchmarking On-Device Models as Phone Secretaries

Unknown authors
· 0 citations · 73 references

TL;DR

CallScreenBench is presented, which scores this setting on five quality dimensions, each printed beside the counter-metric that bills it and never averaged into one number, alongside a guard-edness profile for a toolless proxy that holds no credentials and calls no tools.

View source

Similar papers

#natural language process... Preprint Sep 2026

Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model

Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, one calibrated probability per option comes out of a single forward pass, with no generated text. We test an open implementation of this read...

Si-Miao Ren, Kidus Zewde, Xing-Yu Shen et al. · 6 citations · ⚡1
#artificial intelligence Preprint Sep 2026

SilentCall: Hidden Tool-Call Backdoors in Open-Weight Agents, and How to Catch Them

Open-weight tool-calling agents are adopted on evidence of merit, usually benchmark scores and a record of reliable use. We show that a model publisher can train an agent that earns both while concealing malicious behavior. Fine-tuned on a mixture of clean and poisoned conversations, our agents answer ordinary requests...

Bhanu Pallakonda, Mikkel Hindsbo, Sina Ehsani et al. · 0 citations
#machine learning Preprint Sep 2026

Backdoor Mitigation in Decentralized LLM Fine-Tuning

Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This setting, however, is v...

Sayan Biswas, Jade Garcia Bourrée, R. Guerraoui et al. · 0 citations
#artificial intelligence Preprint Aug 2026

String: An Agentic OS Where Every App Is a Markdown File

String is presented, an open-source runtime that gives this new class of software user an interface of its own and treats the job as an operating-systems problem and what three months of production use taught us.

Jookyung Song, Nojun Kwak, Simyung Chang · 0 citations
Preprint Aug 2026

"Operator, can you hear me?"A Faithful Line into the UNISOC Baseband

Baseband processors are reachable over the radio at all times. Their most security-relevant logic runs deep inside protocol state machines: the control-plane handlers that gate registration, authentication, and session setup. Analyzing that logic systematically requires introspecting the firmware as it runs, which make...

Eduard Vlad, Philipp Mao, Marcel Busch et al. · 0 citations
#artificial intelligence Review Sep 2026

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations...

Pritish Mishra, Ishaan Kumar, Akshat Mandoli et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.