Skip to content

Similar papers

#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Conference Aug 2026

Evaluating AutoEDA Without Human Raters: A Diagnostic Framework and Case Study of Structured Question Generation

Human evaluation remains the dominant way to assess automated exploratory data analysis (AutoEDA) systems, but it is expensive, subjective, and hard to reproduce. We introduce a reproducible automatic diagnostic framework that surfaces structural and statistical failures in AutoEDA—failure modes often under-emphasized...

B. N. Do, Phu-Vinh Nguyen, Hung-Nghiep Tran · 0 citations
Preprint Aug 2026

Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune

Attune is presented, a mixed-initiative system for steerable LLM-powered scoring that performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process.

Bhavya Chopra, Meng Chen, Rebecca Dang et al. · 0 citations

AI-Powered Resume

A dual-engine, AI-powered resume screening system designed for transparency and reproducibility, with a reproducible SBERT→XGBoost→SHAP classification pipeline, and a practitioner-oriented user interface that operationalizes explainability and auditability is presented.

Sang Suh, Numery Zaber · 0 citations
#natural language process... Preprint Sep 2026

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

NAPHA (eNtropy-Aware Post-Hoc Alignment), a simple yet effective lightweight post-hoc alignment method that matches the LLM distribution to the HJD by first assigning an instance to a discrete entropy class and then routing it to specialized, trained alignment models is proposed.

Sebastian Steindl, Nikos Voskarides, Alberto Gasparin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.