ViG: Grounded Adaptation of Transformer-Based Models for Low-Resource Vietnamese Image Captioning
Abstract
Dual-feature captioners that fuse region-level and grid-level visual features define the state of the art in image captioning, yet they owe this strength to visual encoders trained on massive English corpora. A low-resource language such as Vietnamese cannot retrain these encoders, and finetuning the decoder alone cannot recover what the encoder never learned to see, including how objects interact and the culturally specific content that fills everyday Vietnamese scenes. We propose ViG, which grounds the frozen GRIT captioner in this missing evidence through two lightweight gated-residual branches. A Residual Relation Memory enriches each detected region with pairwise geometric and appearance cues, recovering the actions, spatial relations, and counts that isolated regions cannot express. A Local-Cultural Query Memory distills Vietnamese-specific visual patterns into a compact set of learned queries, guided by phrases mined from the training captions. Both branches open only where the Vietnamese data provides signal, preserving the pretrained captioner they build on. ViG improves over the GRIT baseline on every standard metric on both KTVIC and UIT-ViIC, and outperforms a substantially larger fine-tuned multimodal model. Grounding a pretrained captioner in targeted visual evidence thus emerges as a promising direction for image captioning in low-resource languages.