Beyond OCR Accuracy: Text-Centric VQA Under Image Degradation with Modular and End-to-End
An empirical study comparing two modular pipelines with SA-DBNet, a custom detector architecture combining ResNet-18 with self-attention spatial modeling and deformable convolutions against an end-to-end vision-language baseline, evaluated on 4013 degraded images with 7000 question-answer pairs finds that conventional...