Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe, however, matches th...
Zhu Liu, Zi-Yi Wang, Yao Zhang et al.· 0 citations
This work presents DistilVDR, a 524M end-to-end VDR system distilled bilaterally from a single 8B vision-language teacher under a pointwise cosine alignment loss and matches VDR's text-query and image-document input asymmetry with an asymmetric encoder-only student that concentrates visual capacity on the document side...
The authors' layer-resolved SR analysis reveals an Alignment-Aggregation Divergence: visual structure is preserved as a stable ``Structural Plateau''within the backbone, while the final layers reshape it into a sparse, query-aligned form unsuitable for pruning.
Zhu Liu, Ziyu Hu, Yao Zhang et al.· 7 citations· ⚡2
NanoVDR exploits query--document asymmetry by decoupling the two encoding paths: a frozen 2B VLM teacher indexes documents offline, while a distilled text-only student as small as 69M parameters encodes queries at inference, and the resulting NanoVDR-S-Multi (DistilBERT, 69M) retains 95.1% of teacher quality.
Zhu Liu, Yao Zhang, Yuntian Xiao· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.