Preprint
Aug 2026
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
It is shown that paragraph supervision enables effective use of long token sequences, whereas caption-only training degrades beyond 60 tokens, and paragraph supervision consistently benefits long-description benchmarks and hard negatives prove detrimental in text-only fine-tuning.
Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz et al.
· 0 citations