Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain...
Adam Gaber, Uriel Dolev, Elisabeth Fittschen et al.· 0 citations
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulne...
Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.· 0 citations
This work systematically investigate the effect of anchor selection by evaluating 22 different anchors on the Arena-Hard-v2.0 dataset, finding that the choice of anchor is critical: a poor anchor can dramatically reduce correlation with human rankings.
Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen et al.· Annual Meeting of the Associ...· 6 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.