This work introduces Pluralis v0.1, a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective and calls upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.
Alicia Parrish, Rajat C. Shinde, Sanket Badhe et al.· arXiv.org· 0 citations
A novel dataset of 70 underexplored graph construction problems that require finding Ramsey-good graphs with special properties, which shows that LLMs achieve only 37.70% accuracy on the hard-tier problems in this dataset, with Gemma-4-31B achieving the highest performance out of the five.
This thesis proposal presents a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment, and highlights the importance of human feedback for ensuring that AI systems align with human values.
Deepak Pandita, Christopher Homan· Annual Meeting of the Associ...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.