This work proposes Feedback-to-Rubrics, a problem setting for learning criteria from inline comments on artifacts, which infers rubrics from these comments and iteratively refines them by observing errors in comment prediction based on the inferred rubrics.
Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...
Zi-Hao Hu, F. Fukumoto, Jian He et al.· Scientometrics· 0 citations
As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based met...
Xu-Kai Wang, Liangqi Li, Zhi-Yu Xu et al.· 0 citations
GenRubric is introduced, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution, and experiments show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-wri...
Yifan Chen, Hai-Tao Li, Qing-Yao Ai et al.· 1 citation
This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.
M. Nadăş· Artificial Intelligence Revi...· 0 citations