Preprint
Aug 2026
CallScreenBench: Benchmarking Small Language Models as Phone Secretaries
This work presents CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement, and reports quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.
Jia-Qi Gan, Hao Tang, Jamey Z. Liang et al.
· 1 citation