An empirical study comparing two modular pipelines with SA-DBNet, a custom detector architecture combining ResNet-18 with self-attention spatial modeling and deformable convolutions against an end-to-end vision-language baseline, evaluated on 4013 degraded images with 7000 question-answer pairs finds that conventional OCR error metrics are unreliable predictors of downstream VQA performance.
Abstract
Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a task made substantially harder when images suffer from real-world degradation such as motion blur, low resolution, or compression artifacts. While modular OCR-based pipelines and end-to-end vision-language models are both widely used for this task, their comparative robustness under degraded conditions remains underexplored. We present an empirical study comparing two modular pipelines with SA-DBNet, a custom detector architecture combining ResNet-18 with self-attention spatial modeling and deformable convolutions against an end-to-end vision-language baseline, evaluated on 4013 degraded images with 7000 question-answer pairs. Fine-tuned modular pipelines achieve up to 57.50% exact-match accuracy versus 38.00% for the end-to-end baseline, with domain-specific fine-tuning yielding a gain of up to 29.50 percentage points. Critically, we find that conventional OCR error metrics like Character Error Rate and Word Error Rate are unreliable predictors of downstream VQA performance, as semantic reasoning can compensate for recognition failures when contextual cues are present. These findings highlight the importance of task-aware evaluation for text-centric VQA systems under realistic visual conditions. Codes are available here
It is shown that paragraph supervision enables effective use of long token sequences, whereas caption-only training degrades beyond 60 tokens, and paragraph supervision consistently benefits long-description benchmarks and hard negatives prove detrimental in text-only fine-tuning.
Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz et al.· 0 citations
It is found that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training, highlighting the value of screening backbones with Sink Strength before VL training and narrow the int...
Minsik Choi, Geewook Kim, Young Geun Kim· 0 citations
This work fuse DINOv3 and CleanDIFT representations into a perception encoder and align it with RoBERTa-L, yielding a text-aligned model TDDN that preserves this perceptual advantage and leads on segmentation benchmarks among general-purpose contrastive encoders, including SigLIP.
Harsha Patnala, Debopriyo Banerjee, A. Munot et al.· 0 citations
Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In applications such as product labeling, signage, and interface design, target text should be rendered within a designated text-bearing region w...
Yan Wang, Xin-Yi Hou, Wei-Guo Lin et al.· 0 citations
VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.
Experiments on a new approach for fine-grained evaluation demonstrate that this approach enhances a model’s ability to understand fine-grained differences.
Aozhu Chen, Hazel Doughty, Xirong Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.