Skip to content
Review

Feedback-to-Rubrics: Can We Extract Expert Criteria from Inline Comments?

· 0 citations · 45 references

TL;DR

This work proposes Feedback-to-Rubrics, a problem setting for learning criteria from inline comments on artifacts, which infers rubrics from these comments and iteratively refines them by observing errors in comment prediction based on the inferred rubrics.

View source

Similar papers

Review Open access Aug 2026

LLM aspect prediction: reviewing academic papers from different aspects with Large Language Model

Peer review is a fundamental process in scholarly publishing, wherein reviewers assess and score various aspects of a manuscript (e.g., novelty, clarity, and significance) based on established evaluation criteria. However, this process demands substantial time and effort, and remains inherently susceptible to human bia...

Zi-Hao Hu, F. Fukumoto, Jian He et al. · 0 citations
Preprint Aug 2026

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based met...

Xu-Kai Wang, Liangqi Li, Zhi-Yu Xu et al. · 0 citations
#natural language process... Preprint Aug 2026

GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation

GenRubric is introduced, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution, and experiments show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-wri...

Yifan Chen, Hai-Tao Li, Qing-Yao Ai et al. · 1 citation
Book Open access Aug 2026

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.

Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al. · 0 citations
Review Open access Aug 2026

Large language models as judges: recent advances in LLM-based evaluation, critique, preference modeling, and feedback for text and code

This survey provides a comprehensive overview of recent advances in LLM-based evaluation, covering techniques, applications, and challenges across domains, with future directions emphasizing standardized protocols, uncertainty estimation, and human–AI collaboration.

M. Nadăş · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.