A MULTIMODAL PIPELINE BRIDGING CAPTIONING AND OPEN-VOCABULARY DETECTION FOR ENHANCED VISIONLANGUAGE UNDERSTANDING
This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding. The study evaluates several captioning models, including BLIP, BLIP-2, InstructBLIP, and LLaVA, in combination with two open-vocabulary detectors, OWLv2 and Grounding DINO. Experiments conducted on a representative subset of the COCO dataset show that prompt quality significantly influences detection performance and that post-processing operations, including label normalization and filtering, substantially improve semantic detection metrics. The results reveal complementary behaviors between Grounding DINO and OWLv2, highlighting the importance of prompt engineering and output refinement in multimodal vision-language pipelines. Rather than introducing a new detection architecture, this work provides a comparative analysis of the interactions between caption generation, prompt extraction, and open-vocabulary detection, offering insights for the design of future interactive vision-language systems.