Multimodal Context-Enriched Visual Representation Learning for Enhanced Vision–Language Image Captioning
Image captioning models can produce rapid Sentences, without visual relationships, or insert non-existing plausible objects. A common cause is to compress image evidence into visual symbols that carry a weak neighbourhood context. The multimodal context-enhanced visual representation learning framework (MCVRL) addresse...