Skip to content
Book Open access

TEFD: A Benchmark for Natural Language to Flux Query Generation in Time-Series Databases

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 10020-10031 · 0 citations · 30 references

TL;DR

This work introduces the Text-to-Flux task and proposes FluxEngine, a novel automated pipeline for dataset construction, and provides the essential infrastructure to foster future research in NLI for TSDBs.

Abstract

The proliferation of IoT and real-time monitoring has established Time-Series Databases (TSDBs) like InfluxDB as critical infrastructure. However, their functional query languages (e.g., Flux) present a steep learning curve, hindering data accessibility for non-experts. While Natural Language Interfaces (NLIs) offer a potential solution, the domain of Text-to-Flux is stalled by a critical bottleneck: the complete absence of diverse, high-quality paired benchmarks. To address this, we introduce the Text-to-Flux task and propose FluxEngine, a novel automated pipeline for dataset construction. Unlike static generation methods used in Text-to-SQL, our framework features a Self-Sustaining Live Data Context that utilizes background tasks to perpetually generate fresh data, ensuring that queries involving relative time windows (e.g., ''past hour'') remain executable and valid indefinitely. Using this framework, we construct and release TEFD (Text-to-Flux Dataset), the first large-scale benchmark for this task. We further define execution-based evaluation metrics tailored for time-series validity. This work provides the essential infrastructure to foster future research in NLI for TSDBs. To facilitate reproducibility and future research, our dataset and benchmark code are publicly available at https://github.com/gta886/TEFD-Benchmark.

Read PDF

Similar papers

Book Open access Aug 2026

Automating End-to-End Hybrid Query Processing: Benchmark, Solution, and Insights

Hybrid queries—natural language questions over structured data that require both database capabilities and LLM reasoning—have recently emerged as a prominent research topic. However, existing solutions remain overly dependent on manual workflows, and current benchmarks are limited in scale and diversity. To bridge this...

Bo Li, Chenzhan Wang, Long-Kang Lin et al. · 0 citations
Conference Aug 2026

A Template-Driven Multi-Source Benchmark Generation Method for Heterogeneous Data Lakes

Large language models (LLMs) have rapidly advanced natural-language-to-query (Text-to-Query) capabilities, yet existing public benchmarks remain confined to single database paradigms such as Text-to-SQL or Text-to-KG. They do not capture real-world settings where relational databases, graph databases, document database...

Guo-Shen Li, Hang Zhang, Ying-Jun Liu et al. · 0 citations
Conference Aug 2026

"Question→SQL→Wiki" Dynamic Wiki Graph for NL2SQL

To bridge the semantic gap in NL2SQL (Natural Language to SQL) tasks, this study proposes a "Question→SQL→Wiki" framework that leverages a dynamic Wiki Graph as an intermediate reasoning layer. Departing from conventional NL2SQL approaches that rely solely on end-to-end mapping, our method utilizes Large Language Model...

Jia-Xuan Liu, Shi-Yu Fang, Ji-Bing Wu et al. · 0 citations

TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases

Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-s...

Fei Lyu, Zhi-Yi Peng, Jia-Ming Liu et al. · 0 citations
Preprint Aug 2026

Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

DBLifeBench is introduced, the first benchmark to evaluate LLMs across five critical lifecycle phases: Design, Implementation, Operation, Debugging, and Maintenance, and a novel task utilizing structured reasoning graphs to mimic human iterative problem-solving is proposed.

Shunfan Zheng, Dong-Sheng Shi, Yue Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs

Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semantic states change over time. However, existing benchmarks primarily emphasize structural evolution via temporal link prediction (TLP), while support for semantic evolution remains limited. Although temporal node...

Longfei Ma, Ze-Min Liu, Fei Wu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.