Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

A Template-Driven Multi-Source Benchmark Generation Method for Heterogeneous Data Lakes

Large language models (LLMs) have rapidly advanced natural-language-to-query (Text-to-Query) capabilities, yet existing public benchmarks remain confined to single database paradigms such as Text-to-SQL or Text-to-KG. They do not capture real-world settings where relational databases, graph databases, document databases, and NoSQL systems coexist. We propose a template-driven method for constructing a multi-database Text-to-Query benchmark targeting heterogeneous data lakes. Our framework introduces a unified semantic entity space anchored by a Global Entity Registry (GER) that maps local identifiers across SQL, Neo4j, MongoDB, and document retrieval systems. From this representation we design 400 query templates and generate approximately 18380 cross-database instances via template-driven entity sampling. Every instance passes execution-level verification and structural consistency validation. The resulting dataset supports single-database queries, cross-database reasoning, and multi-source fusion tasks, offering a standardized resource for evaluating LLMs under heterogeneous data conditions. The benchmark is built on a fully synthetic hierarchical-organization scenario; the template-driven generation method and GER-based entity alignment are domain-agnostic and directly applicable to any multi-model data environment.

Guo-Shen Li, Hang Zhang, Ying-Jun Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.