A culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios, and a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios to construct culturally situated data are introduced.
Ze-Kun Yuan, Yang-Fan Ye, Bao-Hang Li et al.· 0 citations
This work proposes a training-free, interpretable framework that selects thinking words via attention heads to guide LRMs toward more effective reasoning, and consistently outperforms the strong baseline DEER.
De-Zhi Zhao, Xin Liu, Xiao-Cheng Feng et al.· Proceedings of the Thirty-Fi...· 0 citations
An Anchor-Based Cross-Tokenizer Distillation with Residual Regularization with Residual Regularization (ACTD) bridges structural heterogeneity through vocabulary and sequence alignment, while mitigating alignment noise via a novel anchor loss with residual regularization.
Huiyi Zhang, Zijian Li, Xiaocheng Feng et al.· 0 citations
Experiments on challenging reasoning benchmarks show that H$^2$SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.
Qi Cai, Yi-Chuan Ma, Linyang Li et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.