This work proposes DocPC, a document-level visual retrieval framework based on Representative Page Composition, and introduces DocViRe, a benchmark with multi-positive relevance annotations that combines multi-positive contrastive learning with sparsely scheduled listwise optimization.
Abstract
Visual document retrieval has advanced by encoding page screenshots with vision-language models, bypassing OCR pipelines. However, existing methods remain page-centric, misaligned with real-world scenarios requiring complete document retrieval. A naive page-then-document aggregation suffers from linear indexing cost and degraded retrieval when relevance spans multiple pages. We propose DocPC, a document-level visual retrieval framework based on Representative Page Composition: selecting representative pages and composing them into a single grid image for document-level indexing, reducing indexed images, vectors, and storage by 10.1x and end-to-end indexing time by roughly 7.7x. To handle multi-positive supervision prevalent at the document level, we combine multi-positive contrastive learning with sparsely scheduled listwise optimization. We also introduce DocViRe, a benchmark with multi-positive relevance annotations. DocPC-ColQwen achieves NDCG@5 of 44.09 on DocViRe, outperforming the strongest page-level baseline at 38.91 while reducing storage by 10.1x. Code is available at https://anonymous.4open.science/r/DocPC-Document-Level-Visual-Retrieval-via-Representative-Page-Composition-1D52. Data is available at https://huggingface.co/datasets/anonymous-7219/docpc.
ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
Adrien Mialland, Marc Plantevit, Julien Gallois et al.· 0 citations
MIDR (Multimodal Indexing for Document Retrieval), a training-free framework for enrichment-augmented indexing that shifts multimodal reasoning to index time, establishes index-time multimodal reasoning as a compelling accuracy-deployment alternative to serving-time visual late interaction.
Debanjan Mahata, Atharva Tendle, Daniel Preoţiuc-Pietro et al.· 0 citations
It is demonstrated that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retr...
Scene Text Retrieval (STR) aims to search images containing a given textual query within large-scale image collections. However, existing approaches are fundamentally constrained in two ways: 1) they are evaluated on narrow benchmarks that focus primarily on natural scenes; and 2) they rely on either error-prone multi-...
Tong-Kun Guan, Yu-Tong Cai, Hao-Cheng Wang et al.· IEEE Transactions on Image P...· 0 citations
It is argued that effective compression should preserve query-relevant coverage, meaning the diverse document regions that may become the strongest MaxSim match across queries, rather than selecting patches independently by salience, why dense rendered pages are easier to compress than natural images.
Ailar Mahdizadeh, Aria Salari, Sohail Rajabi et al.· 0 citations
Results show that late-interaction visual retrieval benefits from query-aware allocation rather than query-agnostic pooling alone, and identifies token top-k for interactive retrieval and the naive greedy implementation as a quality upper envelope.
PS Rishi, R. Dwivedi, V. K. Kurmi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.