Skip to content
Preprint

DocPC: Document-Level Visual Retrieval via Representative Page Composition

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

This work proposes DocPC, a document-level visual retrieval framework based on Representative Page Composition, and introduces DocViRe, a benchmark with multi-positive relevance annotations that combines multi-positive contrastive learning with sparsely scheduled listwise optimization.

Abstract

Visual document retrieval has advanced by encoding page screenshots with vision-language models, bypassing OCR pipelines. However, existing methods remain page-centric, misaligned with real-world scenarios requiring complete document retrieval. A naive page-then-document aggregation suffers from linear indexing cost and degraded retrieval when relevance spans multiple pages. We propose DocPC, a document-level visual retrieval framework based on Representative Page Composition: selecting representative pages and composing them into a single grid image for document-level indexing, reducing indexed images, vectors, and storage by 10.1x and end-to-end indexing time by roughly 7.7x. To handle multi-positive supervision prevalent at the document level, we combine multi-positive contrastive learning with sparsely scheduled listwise optimization. We also introduce DocViRe, a benchmark with multi-positive relevance annotations. DocPC-ColQwen achieves NDCG@5 of 44.09 on DocViRe, outperforming the strongest page-level baseline at 38.91 while reducing storage by 10.1x. Code is available at https://anonymous.4open.science/r/DocPC-Document-Level-Visual-Retrieval-via-Representative-Page-Composition-1D52. Data is available at https://huggingface.co/datasets/anonymous-7219/docpc.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering

ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-$k$ retrieval method for late-interaction visual document retrieval, is introduced and it is shown that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.

Adrien Mialland, Marc Plantevit, Julien Gallois et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval

MIDR (Multimodal Indexing for Document Retrieval), a training-free framework for enrichment-augmented indexing that shifts multimodal reasoning to index time, establishes index-time multimodal reasoning as a compelling accuracy-deployment alternative to serving-time visual late interaction.

Debanjan Mahata, Atharva Tendle, Daniel Preoţiuc-Pietro et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval

It is demonstrated that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retr...

Trishan Singha Roy, Arkadeep Acharya, Vishwajeet Kumar et al. · 0 citations
Sep 2026

BPE-Level Visual-Textual Alignment for Multi-Scene Text Retrieval

Scene Text Retrieval (STR) aims to search images containing a given textual query within large-scale image collections. However, existing approaches are fundamentally constrained in two ways: 1) they are evaluated on narrow benchmarks that focus primarily on natural scenes; and 2) they rely on either error-prone multi-...

Tong-Kun Guan, Yu-Tong Cai, Hao-Cheng Wang et al. · 0 citations
Preprint Aug 2026

Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers

It is argued that effective compression should preserve query-relevant coverage, meaning the diverse document regions that may become the strongest MaxSim match across queries, rather than selecting patches independently by salience, why dense rendered pages are easier to compress than natural images.

Ailar Mahdizadeh, Aria Salari, Sohail Rajabi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.