This work proposes MetaPersona, a framework that retrieves task-relevant evidence, constructs literature-derived persona dependency graphs, and samples synthetic populations from empirical priors linking demographics, latent attributes, and outcomes.
Jin Ye, Yuan-Gang Li, Chen-Xiao Yu et al.· 0 citations
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating...
Cristina V. Lopes, Yuan-Gang Li, Justin Tian Jin Chen et al.· 0 citations
This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.
Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.