Global-to-Local Visual Conditioning for Image Captioning with a Frozen Vision–Language Model
Image captioning with pretrained vision and language models often requires substantial model adaptation, while lightweight settings restrict the number of components that can be updated. This study investigates an input-level visual conditioning approach for image captioning using a frozen ResNet-50 image encoder and a...