Skip to content
Preprint

When Is Complex Chunking Worth It? A Multi-Objective Evaluation of Chunking Methods at Scale

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

This work evaluates eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs and shows that computationally expensive methods rarely provide consistent gains over simpler chunking.

Abstract

Dense retrieval is commonly evaluated on benchmarks that represent each document with a single embedding, even though real-world retrieval systems often index long documents that require chunking. In these settings, the chosen chunking method not only affects retrieval quality, but also indexing throughput, query latency, and memory usage. Prior comparisons of chunking strategies have mainly focused on retrieval performance, leaving operational trade-offs underexplored. To address these issues, we evaluate eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs. Our results show that computationally expensive methods rarely provide consistent gains over simpler chunking. Instead, the best performing strategy depends on the embedding model, dataset, corpus size, and target retrieval metric. Methods with similar performance can also differ substantially in operational cost, showing that chunking should be seen as a multi-objective design decision.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.

Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak · 1 citation
Open access 2026

Effective Chunking for Retrieval-Augmented Generation over Structured Institutional Documents

Retrieval-Augmented Generation (RAG) has improved domain-specific question answering tasks by grounding responses in an external knowledge base. However, institutional documents differ from general corpora due to the presence of structured content such as tables, flowcharts, fee breakdowns and program codes. Existing R...

Siti Sarah Izhan Khalib, Abdul Hadi Abd Rahman, L. Zakaria · 0 citations
Preprint Sep 2026

A Systematic Multi-Domain Evaluation of Document Retrievers

Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks...

Valentin Velev, Andreas Spitz · 0 citations
Open access Sep 2026

Chunk Size and Retrieval Depth Optimization in a Minimal Retrieval-Augmented Generation Pipeline

Retrieval‑augmented generation (RAG) links a large language model to a set of texts. But RAG works well only when two settings are right: how big each piece of text (a "chunk") is, and how many pieces the model reads. People often choose these by habit, not by data. In this study we tested how both settings change the...

Sajjad Ahmad, Jasim Hussain, Hamza Najeeb et al. · 0 citations
Preprint Aug 2026

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the curr...

Lin-Hai Ma, Ethan F. Wei, Xue-Qing Peng et al. · 0 citations
#small language model Preprint Aug 2026

Query Expansion Is More Than Generation: Improving Dense Retrieval through Better Integration

This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...

Sixia Sun, Mihai Surdeanu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.