Skip to content

Diagnostic Knowledge Graphs: Automated Benchmark Construction and Deterministic Evaluation for Multi-Step Reasoning Agents

Unknown authors
· 0 citations · 26 references

TL;DR

It is argued that building reproducible diagnostic benchmarks and building effective diagnostic agents are dual problems solved by the same artifact —a canonical knowledge structure that normalizes evaluation gold-standards and constrains agent hypothesis spaces simultaneously.

View source

Similar papers

Preprint Aug 2026

Evidence-Carrying Validation for Knowledge Graphs

This work presents an evidence-carrying validation interface: every selected node-shape check returns either a satisfaction trace or failure witness, and shows how programs combine passing and failing evidence to diagnose missing information and guide repair.

Gabe Fierro · 0 citations
Preprint Aug 2026

Towards Researcher Agents for Knowledge-Graph Question Answering

This work presents an agentic text-to-SPARQL system that goes one step beyond static tool-using agents: a researcher agent that, after each round of inference on a validation set, proposes and tests changes to its own prompts, rules, and tool-orchestration code.

Tommaso Soru, Abdulsobur Oyewale · 0 citations
#machine learning Preprint Aug 2026

ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning

ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...

S. Goria · 0 citations
#artificial intelligence Preprint Sep 2026

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench is introduced, a controlled multi-turn benchmark for measuring model-level behavior across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.

Yaxing Lyu, Sheng-Jie Zhou, B. Toh et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ASTRA - Agentic System for Ticket Resolution and Analysis

ASTRA, an agentic system for ticket resolution in which a central orchestrator coordinates three specialist information-gathering agents and drives a judge-orchestrator refinement loop to produce evidence-backed troubleshooting reports, is proposed.

Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao et al. · 0 citations
Conference Aug 2026

TAP-LLM: An Executable Attribution Framework for Temporal Event Graph Prediction

Although temporal event graph predictors can infer future relational events from historical sequences, their scores provide limited evidence about which historical events support a particular output. We propose TAP-LLM, an executable attribution framework for temporal event graph prediction. Rather than treating explan...

Wan-Ying Liu, Wen Zhou, Jian-Bo Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.