In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600...
Timothee Mickus, Claudio Savelli, Eduardo Calò et al.· 3 citations
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulne...
Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.· 0 citations
This time, the SHROOM-Visions task aims to tackle hallucinations through a model-agnostic detection task focused on large vision-language models, building on the recently introduced SHEEP dataset.
Raúl Vázquez, Aman Sinha, Chuyuan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.