Skip to content

Low-Latency Semantic Processing with Optimal Prompt-Level Batching

· 0 citations · 14 references

TL;DR

This paper measures and analytically derive an optimal size for in-prompt document batching to effectively amortize this overhead of LLM calls over public APIs, cutting end-to-end semantic operator latency by up to 14 × with no meaningful accuracy loss.

View source

Similar papers

#machine learning Preprint Sep 2026

RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple lo...

Armin Oliya, Aleksandra Sawczuk, Radosław Białobrzeski · 0 citations

Maintainable, Low-Latency and High-Quality Query Compilation with MLIR and TPDE

A low-latency MLIR compilation backend based on TPDE that by-passes the LLVM lowering and compilation pipeline is introduced and techniques that reduce MLIR’s intrinsic overhead are proposed.

Jonas Ladner, Lukas Döllerer, Alexis Engelke et al. · 0 citations
Open access Aug 2026

Real-time NLP moderation on XMPP with GPU-accelerated inference and adaptive batching

Real-time communication platforms generate continuous streams of short, latency-sensitive messages, creating a demanding environment for automated toxicity detection. Transformer-based language models offer strong contextual accuracy, but their inference cost makes them difficult to deploy efficiently in decentralized...

A. Aldarf, A. Shaker, I. Bessmertny · 0 citations
Jul 2026

Accelerating String-Heavy Queries with LLM Token Tables

Strings are the most common data type in modern database systems, yet they are often treated as an afterthought in high-performance data formats. While numerical data benefits from specialized, lightweight compression schemes, text is typically handled by general-purpose algorithms such as Zstd, LZ4, or Snappy, which r...

Tobias Schmidt, Nicolas Schmitt, Thomas Neumann et al. · 0 citations
Preprint Sep 2026

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration

Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining service completion. The in...

De-Long Li, Xu Wang, Hao-Chen Gong et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.