Skip to content
Conference

ViG: Grounded Adaptation of Transformer-Based Models for Low-Resource Vietnamese Image Captioning

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 652-657 · 0 citations · 18 references

Abstract

Dual-feature captioners that fuse region-level and grid-level visual features define the state of the art in image captioning, yet they owe this strength to visual encoders trained on massive English corpora. A low-resource language such as Vietnamese cannot retrain these encoders, and finetuning the decoder alone cannot recover what the encoder never learned to see, including how objects interact and the culturally specific content that fills everyday Vietnamese scenes. We propose ViG, which grounds the frozen GRIT captioner in this missing evidence through two lightweight gated-residual branches. A Residual Relation Memory enriches each detected region with pairwise geometric and appearance cues, recovering the actions, spatial relations, and counts that isolated regions cannot express. A Local-Cultural Query Memory distills Vietnamese-specific visual patterns into a compact set of learned queries, guided by phrases mined from the training captions. Both branches open only where the Vietnamese data provides signal, preserving the pretrained captioner they build on. ViG improves over the GRIT baseline on every standard metric on both KTVIC and UIT-ViIC, and outperforms a substantially larger fine-tuned multimodal model. Grounding a pretrained captioner in targeted visual evidence thus emerges as a promising direction for image captioning in low-resource languages.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.