VAmoS Energy, a benchmark that combines challenges in 100 calls about utility billing and payment assistance, is introduced, showing why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.
Joshua Meyer, S. Shayegan, Ritiz Tambi et al.· 0 citations
Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact c...
Joshua Meyer, S. Shayegan, Ritiz Tambi et al.· arXiv.org· 2 citations
The cluster loop yields the strongest held-out rubric on both evaluator tasks from a commercial search vertical, and is the only method robustly positive on both.
Jinyoung Kim, N. Corp, Sun Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.