Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evalua...
Bryan Chen Zhengyu Tan, Wei-Hua Zheng, Thong T. Doan et al.· 0 citations
Across seven Southeast Asian languages, broad capability gains are observed with the 120B-A12B model showing broader and more consistent improvements across tasks, with the 30B-A3B model showing broader and more consistent improvements across tasks.
Adila Aulia, Ahmed Dabeer, Ahn Jeongmi et al.· 0 citations
CultureConverse is introduced, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains and performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-...
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.