Image captioning using a transformer with topic–word semantic modeling and multimodal feature fusion
Despite the recent advances in Transformer-based image captioning models, reliance on implicit semantic representations and the lack of integration of topic-level and word-level semantic information with appearance and geometric features remain challenging. To address these limitations, we propose a semantic modeling f...