In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make benchmarking detection independent of particular models. To this end, we construct a dataset of 1,600...
Timothee Mickus, Claudio Savelli, Eduardo Calò et al.· 3 citations
This work evaluates representative KD methods both on bespoke MT models and LLMs, by considering both translation quality and computational cost, using the Machine Learning Life Cycle Assessment tool, which accounts for costs throughout the KD model life cycle.
Joseph Attieh, Timothee Mickus, Anne-Laure Ligozat et al.· 0 citations
This time, the SHROOM-Visions task aims to tackle hallucinations through a model-agnostic detection task focused on large vision-language models, building on the recently introduced SHEEP dataset.
Raúl Vázquez, Aman Sinha, Chuyuan Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.