Skip to content

Kalypso: Relational LLM Serving

Jul 2026 · arXiv.org · Vol abs/2607.23815 · 0 citations · 34 references
Computer Science

TL;DR

Kalypso is presented, a relational LLM serving system that exposes an API for semantic query plans and executes them using an adaptive, memory-aware scheduling algorithm, demonstrating that query-aware LLM serving can substantially improve the efficiency of semantic query execution.

Abstract

Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructured data. Existing semantic query processing systems invoke request-centric LLM serving systems that are unaware of the query plan, leaving substantial performance opportunities unused. This paper introduces relational LLM serving, an abstraction that makes LLM serving aware of semantic query structure while preserving query semantics and output accuracy. The key opportunity is pipelined execution across semantic operators: when intermediate tuples flow directly from one operator to the next, their KV-cache state can be reused instead of recomputed. We present Kalypso, a relational LLM serving system that exposes an API for semantic query plans and executes them using an adaptive, memory-aware scheduling algorithm. Kalypso addresses a new online scheduling problem in which pipelined operator execution is coupled with GPU memory pressure management to reuse KV-cache state in the serving engine before eviction. Its scheduler continuously adjusts memory allocations to balance upstream parallelism, downstream progress, and GPU utilization. Our evaluation shows that Kalypso improves query completion time over baselines using request-centric LLM serving, with speedups up to 4.57x across diverse workloads, demonstrating that query-aware LLM serving can substantially improve the efficiency of semantic query execution.

View source

Similar papers

Jul 2026

InferScale: GPU-Native KV Injection for Personalized LLM Serving

This work presents InferScale, a GPU-native LLM memory system that replaces repeated prompt prefilling with reusable KV state, and encodes each memory fact together with a small window of preceding conversation context while caching only the target fact's KV.

Peter Li, Prashant Pandey · 1 citation
Preprint Aug 2026

DAGSmith: Dependency-Aware Rewriting for dbt-Style SQL Pipelines

DAGSmith is introduced, to the best of the authors' knowledge the first holistic dependency-aware source-to-source rewriting system for SQL pipeline DAGs and enables dependency-edge simplification, non-local semantic reuse, downstream-aware pruning, pipeline-aware work placement, rewrite-materialization co-optimization...

Jie Liu, Lin Ma, Barzan Mozafari · 0 citations
Jul 2026

The Data World is Not Flat: Efficient Factorized Execution for Relational Systems

A novel code-generating engine with factorization that enables intra-query-parallelized query execution on factorized representations and generates code to overcome their CPU-unfriendly layout, offering a unified and scalable solution for modern workloads.

Stefan Lehner, Thomas Neumann · 0 citations
Jul 2026

Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads

This work builds on Rhyme, a declarative language whose object-notation syntax mirrors the structure of query results, and refine Rhyme's semantics for generator binding and missing values, allowing co-iteration, inner/outer joins, and nested-loop traversals to be expressed under different uses of generator symbols.

Ran Guo, Tiark Rompf · 0 citations
Open access 2026

LLM-augmented query optimization: a hybrid framework for intelligent SQL performance tuning

LLM-QOpt++ is presented, a novel hybrid, confidence-aware query optimization framework that unifies traditional CBO estimation, machine learning–based cost prediction, and large language model (LLM) reasoning within a single adaptive pipeline.

Hanan Abed Alwally Abed Allah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.