This work evaluates eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs and shows that computationally expensive methods rarely provide consistent gains over simpler chunking.
Abstract
Dense retrieval is commonly evaluated on benchmarks that represent each document with a single embedding, even though real-world retrieval systems often index long documents that require chunking. In these settings, the chosen chunking method not only affects retrieval quality, but also indexing throughput, query latency, and memory usage. Prior comparisons of chunking strategies have mainly focused on retrieval performance, leaving operational trade-offs underexplored. To address these issues, we evaluate eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs. Our results show that computationally expensive methods rarely provide consistent gains over simpler chunking. Instead, the best performing strategy depends on the embedding model, dataset, corpus size, and target retrieval metric. Methods with similar performance can also differ substantially in operational cost, showing that chunking should be seen as a multi-objective design decision.
This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.
Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak· 1 citation
Retrieval-Augmented Generation (RAG) has improved domain-specific question answering tasks by grounding responses in an external knowledge base. However, institutional documents differ from general corpora due to the presence of structured content such as tables, flowcharts, fee breakdowns and program codes. Existing R...
Siti Sarah Izhan Khalib, Abdul Hadi Abd Rahman, L. Zakaria· International Journal of Adv...· 0 citations
Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks...
Retrieval‑augmented generation (RAG) links a large language model to a set of texts. But RAG works well only when two settings are right: how big each piece of text (a "chunk") is, and how many pieces the model reads. People often choose these by habit, not by data. In this study we tested how both settings change the...
Sajjad Ahmad, Jasim Hussain, Hamza Najeeb et al.· International Journal of Inn...· 0 citations
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the curr...
Lin-Hai Ma, Ethan F. Wei, Xue-Qing Peng et al.· 0 citations
This work introduces AnchorQE, a training-free method that separately encodes the original query and its expansion before interpolating them, and shows that AnchorQE improves retrieval effectiveness by up to 12.89% when compared to widely-used expansion-only or text-level concatenation baselines across TREC-DL, LoTTE,...
Sixia Sun, Mihai Surdeanu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.