Findings show that the proposed deterministic-first architecture can preserve an auditable structural baseline, quantify observed primary-key and relationship gaps, and prevent model-generated hypotheses from being silently promoted to source-grounded architectural facts.
Abstract
Legacy database migrations often begin with incomplete or outdated documentation, leaving physical data definition language (DDL) as the principal evidence of data architecture. However, DDL does not fully encode conceptual intent, and model-generated completions can be plausible without being correct. This study proposes and evaluates a provenance-aware, deterministic-first pipeline for reconstructing logical and conceptual data specifications from Oracle-oriented DDL while explicitly separating observed facts, deterministic derivations, and large language model (LLM) suggestions. The pipeline performs DDL investigation, parsing, consolidation, primary-key backfilling, type normalization, and declared relationship-graph construction before optional LLM-assisted enrichment. It preserves source provenance in the deterministic catalog and declared relationship graph, and records inferred primary-key and foreign-key candidates in a separate reviewable overlay. We evaluated the implementation on 249 artifactized schema samples comprising 1,225 SQL files. The pipeline completed 244 samples (97.99%); 52 completed samples contained no extractable DDL. Across completed samples, the deterministic path reconstructed 208 tables and recovered 278 declared foreign-key records; 168 of the reconstructed tables lacked an explicitly parsed primary key before backfilling. LLM enrichment generated 100 foreign-key candidates in 36 samples, but the parent-table admissibility rate was only 17.9% for logical-specification candidates and 17.5% for conceptual-specification candidates. These findings show that the proposed deterministic-first architecture can preserve an auditable structural baseline, quantify observed primary-key and relationship gaps, and prevent model-generated hypotheses from being silently promoted to source-grounded architectural facts.
Infrastructure drift can separate a running Kubernetes system from its documented architectural intent. This paper investigates live architecture models: editable architectural representations connected to selected runtime facts through explicit correspondences and recurring conformance checks. The approach is instanti...
Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration and sharing domain, but the shared tier may key entries only by token content and coarse model metadata. This boundary...
Wei Song, Yu-Xin Cao, Xi Zheng et al.· 0 citations
This demo presents SAGE as a pipeline-native runtime system and demonstrates it through an OPC-facing control plane and includes a compact distributed comparison against a LangChain RPC baseline over a shared 15-workload RAG suite, where SAGE nearly doubles full-RAG throughput under matched 8-node settings.
Jun Liu, Shu-Hao Zhang· Workshop Proceedings of the...· 0 citations
Java middleware may expose Java Management Extensions (JMX) through Jolokia’s Hypertext Transfer Protocol (HTTP) bridge. In affected ActiveMQ deployments, reachable Log4j 2 configuration managed beans (MBeans) become write capabilities and, with compatible triggers, enable remote code execution (RCE). We ask: in a spec...
A. Caciulescu, Matei Badanoiu, R. Rughinis et al.· Computers· 0 citations
An automated gap analysis shows the architecture demonstrates alignment with five of eight requirements derived from the EU AI Act and FDA AI/ML guidance, and module-level evaluations of claim extraction, evidence binding, and silent-edit detection are reported together with a measured baseline comparison on retraction...
CPSE is introduced, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations that enables low-resource schema execution while preserving the output structure required downstream.
Zi-Xiao Dong, Wei Yang, Zi-Hao Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.