It is found that although the queries triggering slow planning are largely DBMS-specific, recurring pathologies involving correlated subqueries, CTE expansion, repeated subquery expressions, disjunctive joins, and constant folding affect multiple systems.
Abstract
Query planners are typically expected to produce optimized plans quickly, leading many researchers (including the authors of this paper) and practitioners to design systems that assume query planning is a low-cost operation. Using a lightweight agentic search, we show that this assumption does not always hold. Across seven DBMSes, including four commercial systems, we find at least one query per system that takes more than three minutes to plan. In addition to being slow to plan, such queries risk tying up database resources without performing useful work, creating a potential denial-of-service vector. We analyze the queries our search uncovers and compare how the seven systems respond to each pattern. We find that although the queries triggering slow planning are largely DBMS-specific, recurring pathologies involving correlated subqueries, CTE expansion, repeated subquery expressions, disjunctive joins, and constant folding affect multiple systems. We release our uncovered queries along with a curated suite of parameterized query pathologies that researchers and database engineers can use to test planner robustness. Overall, our results show that query planning cannot always be treated as a predictably inexpensive operation and that its latency and robustness deserve further attention from both database researchers and engineers.
Slow queries frequently cause severe performance bottlenecks in database management systems. Diagnosing their root causes online risks exacerbating resource contention, while data privacy regulations often prohibit copying production data to test environments. Synthesizing a proxy database from non-intrusive metadata t...
Zhao-Yang Zhang, Shuang Liu, Deng-Feng Xu et al.· 0 citations
Selection processes, e.g., determining qualified job candidates or picking vendors, can naturally be modeled as relational queries. To ensure that such a query produces legally compliant and ethically sound results, the outputs of the query are often subject to additional constraints including fairness ratios, budget...
Vaishnavi Deshpande, Seok-Gyun Lee, Shatha Algarni et al.· Proceedings of the VLDB Endo...· 0 citations
LLM-based query reformulation can improve retrieval, but no single reformulation strategy is consistently optimal across queries, domains, retrievers, or model backbones. This creates an inference-time decision problem: ``Given an original query and a pool of candidate reformulations, which one should be issued to the...
Hai-Son Le, Negar Arabzadeh, Amin Bigdeli et al.· 0 citations
Prescriptive analytics workloads often require solving package queries over data that changes continuously. A package query (PQ) returns a multiset of tuples satisfying global constraints and optimizing a given objective, a natural formulation of constrained optimization within a database. When each attribute in the...
Vasileios Vittis, Azza Abouzied, Peter J. Haas et al.· Proceedings of the VLDB Endo...· 0 citations
Database Management Systems (DBMSs) support multiple SQL mechanisms for representing intermediate query results, including VIEWs, Common Table Expressions (CTEs), and Temporary Tables (TEMPTs). When these mechanisms are used to represent the same intermediate query result, the corresponding queries are expected to prod...
Xiao-Xu Niu, Gong Chen, Jin-Fu Chen et al.· 0 citations
This work systematically analyze how retriever and generator complexity interacts across factoid and multi-hop question answering (QA), including bridge and composition reasoning tasks, and introduces DRAG, a query-adaptive framework for selecting retriever-generator configurations.
Neeraj Anand, Payel Santra, Partha Basuchowdhuri et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.