This work evaluates eight representative chunking strategies across two scalable corpora, three embedding models, and multiple corpus sizes, measuring both retrieval effectiveness and system-level costs and shows that computationally expensive methods rarely provide consistent gains over simpler chunking.
Laura Caspari, K. G. Dastidar, M. Dinzinger et al.· 0 citations
Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.
S. Zerhoudi, Jelena Mitrović, M. Granitzer· 0 citations
This work extends the QPP paradigm by studying query performance degradation under corpus inflation in dense retrieval systems and proposes simple adaptations to established QPP measures, most notably a top-k vs background Wasserstein distance measure, which yield more consistent associations with degradation and outpe...
Kanishka Ghosh Dastidar, M. Dinzinger, Laura Caspari et al.· Annual International ACM SIG...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.