Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution is available, the validity of an earlier step may become clearer only through its downstre...
Yi-Ming Feng, Naihao Deng, Yu-Long Chen et al.· 0 citations
Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation across major model families. We argue that these benchmarks are too easy to support their rol...
Naihao Deng, Samee Arif, Shuai-Chen Chang et al.· 0 citations
AtlasNLP is introduced, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced, showing that dataset coverage is highly uneven across countries and tasks and language coverage does not imply geographic rep...
Joan Nwatu, Tsedeniya Solomon Amare, Longju Bai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.