Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by...
Jing-Yu Jia, Antoine Remond-Tiedrez, Aaron Alvarez et al.· 0 citations
This work introduces RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem and shows that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators.
Yan-Xuan Yu, Dong Liu, Shu Wang et al.· arXiv.org· 1 citation
This article provides a unified view of context-adaptive inference across three traditions that are usually treated separately: explicit adaptation in statistics, rapid task-specific adaptation in meta-learning and transfer, and implicit adaptation in large foundation models via prompting, retrieval, and expert routing...
Yueting Yao, Caleb N. Ellington, Jingyun Jia et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.