It is argued that building reproducible diagnostic benchmarks and building effective diagnostic agents are dual problems solved by the same artifact —a canonical knowledge structure that normalizes evaluation gold-standards and constrains agent hypothesis spaces simultaneously.
This work presents an evidence-carrying validation interface: every selected node-shape check returns either a satisfaction trace or failure witness, and shows how programs combine passing and failing evidence to diagnose missing information and guide repair.
This work presents an agentic text-to-SPARQL system that goes one step beyond static tool-using agents: a researcher agent that, after each round of inference on a validation set, proposes and tests changes to its own prompts, rules, and tool-orchestration code.
ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...
KC-Bench is introduced, a controlled multi-turn benchmark for measuring model-level behavior across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Yaxing Lyu, Sheng-Jie Zhou, B. Toh et al.· 0 citations
ASTRA, an agentic system for ticket resolution in which a central orchestrator coordinates three specialist information-gathering agents and drives a judge-orchestrator refinement loop to produce evidence-backed troubleshooting reports, is proposed.
Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao et al.· 0 citations
Although temporal event graph predictors can infer future relational events from historical sequences, their scores provide limited evidence about which historical events support a particular output. We propose TAP-LLM, an executable attribution framework for temporal event graph prediction. Rather than treating explan...
Wan-Ying Liu, Wen Zhou, Jian-Bo Yuan et al.· 2026 12th International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.