Skip to content
Preprint

Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

Aug 2026 · 1 citation · 32 references
Computer Science

TL;DR

Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean, and Rosetta addresses the problem they pose first: recovering what columns and values mean from the data itself.

Abstract

Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers, partial or absent documentation. We address the problem they pose first: recovering what columns and values mean from the data itself. Rosetta places a language model inside a verification harness: a deterministic profiler extracts structural evidence (value fingerprints, a 26-pattern library, checksum verdicts), the model proposes semantics conditioned on that evidence, and every fact carries provenance and a confidence bounded by its evidence class. Against human documentation on 680 paired columns across eleven BIRD databases, identifiers destroyed, the harness delivers metadata that is 0.475 accurate on the 42% of columns it commits to, against 0.223 on 94% for the same model used directly. Restricted to the 283 columns where both arms speak, the harness writes no better prose than the model alone; the gain is selection: deterministic evidence governs whether the system speaks (coverage +0.257 [0.128, 0.378]), not how well. The deterministic layer is a competence detector, not a competence amplifier. A backbone swap bounds the claim: the prose finding reproduces, but prompt-requested abstention does not transfer; a code-enforced commit gate (predictions registered first; measured on a third backbone and held-out databases) makes no-evidence coverage 0.000 on every backbone. On a blind i2b2 clinical warehouse Rosetta decodes 95.5% of 134 real ICD-9 codes from values alone and abstains on all 44 NDC drug codes. The catalog supports calibrated abstention at query time: under full schema opacity a naive translator falls from 0.92 to 0.42 execution accuracy while our gate answers at 86% accuracy over 59% coverage. Negative results are reported plainly, including that our own authority ladder is not the mechanism behind the headline.

View source

Similar papers

Review Sep 2026

Specification-Driven Data Architecture Reconstruction: From Physical Code to Logical and Conceptual Specifications

Findings show that the proposed deterministic-first architecture can preserve an auditable structural baseline, quantify observed primary-key and relationship gaps, and prevent model-generated hypotheses from being silently promoted to source-grounded architectural facts.

Oleg Grynets, Olena Pochernina, V. Lyashkevych · 0 citations
Preprint Aug 2026

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Doc2DB-Bench is introduced, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells, which provides a testbed for reliable, auditable, and relationally faithful LLM-...

Zhuo-Wen Liang, Zhengxuan Zhang, Jia-Yang Wang et al. · 1 citation
Aug 2026

ATLAS: Adaptive Text-to-SQL with Lifecycle-Aware Self-Maintaining Context

ATLAS addresses Text-to-SQL in production problems by co-locating schema metadata, semantic annotations, and vector embeddings entirely within a single RDBMS with native vector-index support.

Qing Zhang, Shi-Jing Hu, Zhi-Hui Lu · 0 citations
Review Open access Sep 2026

Provenance-Native Audit Infrastructure for LLM-Maintained Wiki Knowledge Systems

An automated gap analysis shows the architecture demonstrates alignment with five of eight requirements derived from the EU AI Act and FDA AI/ML guidance, and module-level evaluations of claim extraction, evidence binding, and silent-edit detection are reported together with a measured baseline comparison on retraction...

Bai-Ling Zhang · 0 citations
Preprint Aug 2026

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

Eigenius is an open-source, typed knowledge-graph DBMS built on a single premise: answering the audit question ("what do you know, and what is your warranty?") requires a unified kernel, and turns data provenance into a structural invariant rather than a property reconstructed across subsystem boundaries.

Hans-Martin Will, Allen L. Brown, M. Fuchs · 0 citations
Preprint Aug 2026

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

In this ablation, retrieved knowledge-base context provides the largest marginal improvement when added to the full oracle graph, and a distillation procedure that turns historical query profiles into reusable SQL reference cards is optimized.

Kate Gwimm, Carson Eisenach · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.