Context-aware image captioning using cosine transformer attention and optimal vision transformer with improved walrus optimization algorithm
Automatic image captioning aims to acquire the learning of the correlation between the visual and the linguistic cognitive domain through the generation of descriptive captions with respect to input images. But it is affected by the semantic distance between image features and text, and problems in modeling fine-grai...