Skip to content

Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads

Jul 2026 · Proceedings of the VLDB Endowment · 0 citations · 41 references

TL;DR

This work builds on Rhyme, a declarative language whose object-notation syntax mirrors the structure of query results, and refine Rhyme's semantics for generator binding and missing values, allowing co-iteration, inner/outer joins, and nested-loop traversals to be expressed under different uses of generator symbols.

Abstract

Modern data processing spans two worlds: flat relational tables, served by decades of database research producing highly optimized query engines, and nested semi-structured data such as JSON, for which expressive query languages exist but compilation and optimization techniques have been applied far less comprehensively. We ask whether a single query language can express both regimes naturally while compiling to efficient native code. We build on Rhyme, a declarative language whose object-notation syntax mirrors the structure of query results, and contribute on three fronts. We refine Rhyme's semantics for generator binding and missing values, allowing co-iteration, inner/outer joins, and nested-loop traversals to be expressed under different uses of generator symbols. We show that Rhyme's prior dependency-driven loop scheduler can generate incorrect code on hierarchical queries, and present a new scheduler based on finer-grained per-statement constraints that ensures correctness. We introduce a gradual type system and a C code generation backend that emits tag-less, statically typed code and specializes data loading and internal data structures for idiomatic SQL patterns. Our system matches state-of-the-art compiled engines on SQL workloads such as TPC-H and outperforms modern JSON-capable databases and DSLs on JSONBench and other hierarchical queries.

View source

Similar papers

Preprint Sep 2026

Compiling Linear Datalog to SQL for Program Analysis

This paper revisits the connection between Datalog and relational databases, advocating recursive SQL as a backend for Datalog evaluation, and presents a compilation framework that translates Datalog programs, particularly those in the Linear Datalog fragment, into equivalent recursive SQL queries.

Amir Shaikhha, Anna Herlihy, Hung Q. Ngo · 0 citations
#artificial intelligence Preprint Sep 2026

DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question...

Yuan Li, Han-Yun Jiang, Guo-Wei Tian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Retrieval-Augmented Generation for Scientific Code Understanding

Large language models have become central to modern coding assistants, but state-of-the-art systems such as Claude Code or Codex rely on very large, cloud-hosted models with significant computational cost and data-privacy implications. This work investigates whether a useful, fully local coding agent can be built aroun...

Aaron Nobile, Andreas Adelmann, Mohsen Sadr · 0 citations
Preprint Sep 2026

Where the LLM Ends and Reliable Decisions Begin

Systems that turn natural-language descriptions of optimization problems into solver-ready code generally use a language model at every stage, including the final translation from a mathematical formulation into executable model-building code. We propose the ANVIL compiler architecture, where we separate these concerns...

Priyadarshan Patil, A. Basu, Vikas Reddy et al. · 0 citations
Preprint Aug 2026

DAGSmith: Dependency-Aware Rewriting for dbt-Style SQL Pipelines

DAGSmith is introduced, to the best of the authors' knowledge the first holistic dependency-aware source-to-source rewriting system for SQL pipeline DAGs and enables dependency-edge simplification, non-local semantic reuse, downstream-aware pruning, pipeline-aware work placement, rewrite-materialization co-optimization...

Jie Liu, Lin Ma, Barzan Mozafari · 0 citations
Jul 2026

The Data World is Not Flat: Efficient Factorized Execution for Relational Systems

A novel code-generating engine with factorization that enables intra-query-parallelized query execution on factorized representations and generates code to overcome their CPU-unfriendly layout, offering a unified and scalable solution for modern workloads.

Stefan Lehner, Thomas Neumann · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.