Skip to content
Open access

Knowledge Graphs vs. SQL over Structured EHR Data

Jul 2026 · Future Internet · Vol 18, pp. 365 · 0 citations · 14 references

TL;DR

Six retrieval configurations that vary along two axes: backend (a property graph database, a relational database and a dense vector index) and interface design (curated domain-specific tool calls, model-generated queries, full-text search, and single-shot dense retrieval) are compared.

Abstract

Clinical question answering over electronic health records (EHRs) increasingly relies on large language model (LLM) agents that retrieve structured patient data through external tools. Published benchmarks, however, evaluate these systems at a single patient-population size, and rarely measure the effect of backend representation from that of the retrieval interface design. This paper compares six retrieval configurations that vary along two axes: backend (a property graph database, a relational database and a dense vector index) and interface design (curated domain-specific tool calls, model-generated queries, full-text search, and single-shot dense retrieval). The evaluation covers a 334-question bank spanning six categories (simple lookup, multi-hop, temporal, cohort, reasoning, and unanswerable), instantiated at three nested population scales: 200, 2000, and 20,000 alive patients from a single Synthea cohort. Four models are compared: Claude Haiku 4.5, Qwen 2.5 72B, Llama 3.1 8B, and Llama 3.3 70B, spanning closed-frontier and open-source alternatives. Curated tool-calling configurations improve accuracy over retrieval-augmented baselines for capable models, but reduce accuracy for a small open-source model due to function-calling protocol failures. We report how accuracy, latency, and cost evolve with each approach, model size, and cohort size, supported by paired statistical tests and confidence intervals. All benchmark components, databases, and evaluation code are publicly available.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

A Living Benchmark for Information Retrieval from Electronic Health Records

Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated,...

J. Cahoon, C. Stanwyck, Sulaiman Somani et al. · 0 citations

MedSQLX: Translation of Medical Queries into UDF-Centric SQL

An agentic framework that leverages Large Language Models (LLMs) for generating UDF-centric queries from natural language task descriptions in the medical domain is presented, demonstrating that structured tool orchestration with verification loops substantially improves generation quality.

Catlynh Nguyen · 0 citations
Aug 2026

CGX: OCR-enhanced knowledge graph retrieval for explainable heart failure analysis

Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting.

Dat T. Nguyen, Ngoc-Anh Thi Le, Binh T. D. Trinh et al. · 0 citations
Open access Sep 2026

Ontology-aware knowledge graph retrieval-augmented generation for clinical decision support

Effectively retrieving and interpreting the vast, diverse, and largely unstructured data contained within electronic health records (EHRs) present significant challenges for clinical decision support systems. Large language models (LLMs), when applied to complex healthcare datasets, frequently exhibit hallucinations, l...

Deepak Panneerselvam, Sasikala E · 0 citations
#artificial intelligence Preprint Sep 2026

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the execut...

Erfan D. Dehkalani, S. Shankaran, Abbot R. Laptook et al. · 0 citations
Open access Sep 2026

Can a General-Purpose Coding Agent Analyze a Production Hospital Data Warehouse?

Background. Health systems answer most questions by having expert analysts hand-write queries against a complex electronic health record data warehouse, a slow, resource-intensive process. Whether an autonomous coding agent can do this accurately is unknown. Methods. In a single-center quality-improvement evaluation, w...

N. Marshall, W. Haberkorn, J. Faulkenberry et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.