Skip to content
Conference

A Template-Driven Multi-Source Benchmark Generation Method for Heterogeneous Data Lakes

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 9-15 · 0 citations · 13 references

Abstract

Large language models (LLMs) have rapidly advanced natural-language-to-query (Text-to-Query) capabilities, yet existing public benchmarks remain confined to single database paradigms such as Text-to-SQL or Text-to-KG. They do not capture real-world settings where relational databases, graph databases, document databases, and NoSQL systems coexist. We propose a template-driven method for constructing a multi-database Text-to-Query benchmark targeting heterogeneous data lakes. Our framework introduces a unified semantic entity space anchored by a Global Entity Registry (GER) that maps local identifiers across SQL, Neo4j, MongoDB, and document retrieval systems. From this representation we design 400 query templates and generate approximately 18380 cross-database instances via template-driven entity sampling. Every instance passes execution-level verification and structural consistency validation. The resulting dataset supports single-database queries, cross-database reasoning, and multi-source fusion tasks, offering a standardized resource for evaluating LLMs under heterogeneous data conditions. The benchmark is built on a fully synthetic hierarchical-organization scenario; the template-driven generation method and GER-based entity alignment are domain-agnostic and directly applicable to any multi-model data environment.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.