Low-Latency Semantic Processing with Optimal Prompt-Level Batching
This paper measures and analytically derive an optimal size for in-prompt document batching to effectively amortize this overhead of LLM calls over public APIs, cutting end-to-end semantic operator latency by up to 14 × with no meaningful accuracy loss.
Jingyi Qu, Samuel Madden, Tianyu Li
· 0 citations