Skip to content

Semantic Data Systems: From Data Management to Data Understanding

Aug 2026 · Proceedings of the VLDB Endowment · 0 citations · 38 references

Abstract

For decades, data management systems have relied on syntactic, structural, and statistical representations of data, leaving semantic understanding largely to human users. This division has become increasingly problematic as users analyze data disconnected from their original collection context, and yet must determine whether datasets are relevant, compatible, and trustworthy. Large language models (LLMs) introduce a new computational capability: the ability to reason about the meaning of data and leverage accumulated background knowledge. I argue that semantic understanding should be viewed as a new systems primitive for data management, enabling a new class of semantic data systems in which understanding becomes a first-class computational capability that complements syntactic and statistical representations of data. Drawing on our recent work spanning dataset discovery, semantic type discovery, schema matching, data harmonization, and agentic workflows, I show that LLMs alone do not solve data management problems: the central challenge is no longer enabling machines to understand data, but designing systems that can represent, maintain, evaluate, and exploit that understanding. These systems should externalize semantic knowledge into reusable artifacts, combine semantic reasoning with algorithms through structured workflows, subject decisions to human validation, and expose specialized capabilities as composable primitives that agents can orchestrate. This perspective suggests a research agenda for semantic data systems that extends beyond data discovery and integration toward a broader transformation of data management.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.