Multimodal Image Description Generation and Alignment Mechanism Integrating Visual Context
Abstract
Image description generation faces the challenge of insufficient visual semantic structuring in cross-linguistic scenarios. Inspired by systemic functional linguistics and cognitive grammar, this paper designs a Perceptual Visual Context Encoder (PVCE), which transforms pixel signals into semantic propositions with image schema structures through spatial topological attention and coordinate attention. A cross-linguistic semantic-visual joint alignment mechanism is constructed, introducing dependency syntactic structure constraints into the contrastive loss to achieve a three-way equivalence mapping between the source language, target language, and visual propositions. Finally, an adaptive generation strategy integrating contrastive decoding is adopted to suppress redundant expressions without visual basis through visual feedback. Experiments on the Multi30K dataset show that the complete model achieves scores of 42.5, 30.7, 122.4, and 24.5 on BLEU-4, METEOR, CIDEr, and SPICE metrics, respectively. Robustness analysis shows that PVCE maintains a high CIDEr score under structurally destructive perturbations.