ViG: Grounded Adaptation of Transformer-Based Models for Low-Resource Vietnamese Image Captioning
Dual-feature captioners that fuse region-level and grid-level visual features define the state of the art in image captioning, yet they owe this strength to visual encoders trained on massive English corpora. A low-resource language such as Vietnamese cannot retrain these encoders, and finetuning the decoder alone cann...