VAmoS Energy, a benchmark that combines challenges in 100 calls about utility billing and payment assistance, is introduced, showing why voice agents need evaluation across the whole call, including what they say, what they change, and how they handle competing speech.
Joshua Meyer, S. Shayegan, Ritiz Tambi et al.· 0 citations
Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact c...
Joshua Meyer, S. Shayegan, Ritiz Tambi et al.· arXiv.org· 2 citations
A variant of FocusAgent significantly reduces the success rate of prompt-injection attacks, including banner and pop-up attacks, while maintaining task success performance in attack-free settings, highlighting that targeted LLM-based retrieval is a practical and robust strategy for building web agents that are efficien...
Imene Kerboua, S. Shayegan, Megh Thakkar et al.· arXiv.org· 15 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.