Skip to content
Conference

Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning

Jul 2026 · 2026 7th International Conference on Smart Systems and Inventive Technology (ICSSIT) · pp. 1816-1821 · 0 citations · 19 references

Abstract

Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresses this error mode by adding local neighbourhood descriptors, global scene tokens, prefix-conditioned visual doors and adaptive contextual corrections be-fore caption decoding. The encoder is trained with cross-entropy and contrast terms for image–text alignment. MSCOCO 2014’s Karpathy test classification results show that BLEU-4, METEOR, CIDEr and SPICE are more powerful captioning bases. The best configuration received a score of 1.352 of the CIDEr compared to 1.308 of BLIP-2 in the same evaluation protocol. The results of the ablation show that the most important contribution is the visual feature enriched by the context, followed by crossmodal gating and adaptive contextual attention. Qualitative examples show that objects with hallucinations are fewer and that spatial relationships are better recovered.

View source

Similar papers

Open access Jul 2026

Global-to-Local Visual Conditioning for Image Captioning with a Frozen Vision–Language Model

Image captioning with pretrained vision and language models often requires substantial model adaptation, while lightweight settings restrict the number of components that can be updated. This study investigates an input-level visual conditioning approach for image captioning using a frozen ResNet-50 image encoder and a...

Sohyun Lee, Dae-Nyoung Heo · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
Aug 2026

Boosting scene captioning with pairwise semantic spatial reasoning and contextualized large language model integration

This work extends its previous work in scene graph generation to integrate image captioning, incorporating large language models to interpret scene graph structures and produce high-quality captions encompassing complex actions and relationships.

Anfel Amirat, N. Baha, Lamine Benrais · 0 citations
Open access Aug 2026

Dual Text-guided Cross-attention for Global–local Visual Fusion in Vietnamese Visual Question Answering

This study proposes a dual-stream architecture that exploits complementary Transformer-based and convolutional visual representations for Vietnamese VQA, and shows improvements over the corresponding single-stream convolutional baselines, supporting the complementary role of the two visual representations.

Huy Tran, V. Nguyen · 0 citations
Open access Jul 2026

A Unified Multimodal Search Framework Using Generative AI and Image Understanding for Enhanced Information Retrieval

A unified multimodal search framework that integrates generative AI-based captioning and image understanding for improved retrieval, enabling more accurate, context-aware search in applications such as e-commerce, multimedia, and large-scale retrieval.

Saeed Alzahrani, Farah Mohammad, Nazar Hussain · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.