Preprint
Aug 2026
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study
The results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil et al.
· 0 citations