Skip to content
Conference

Multimodal Image Description Generation and Alignment Mechanism Integrating Visual Context

Aug 2026 · International Workshop on Artificial Intelligence and Cognition · pp. 963-968 · 0 citations · 15 references

Abstract

Image description generation faces the challenge of insufficient visual semantic structuring in cross-linguistic scenarios. Inspired by systemic functional linguistics and cognitive grammar, this paper designs a Perceptual Visual Context Encoder (PVCE), which transforms pixel signals into semantic propositions with image schema structures through spatial topological attention and coordinate attention. A cross-linguistic semantic-visual joint alignment mechanism is constructed, introducing dependency syntactic structure constraints into the contrastive loss to achieve a three-way equivalence mapping between the source language, target language, and visual propositions. Finally, an adaptive generation strategy integrating contrastive decoding is adopted to suppress redundant expressions without visual basis through visual feedback. Experiments on the Multi30K dataset show that the complete model achieves scores of 42.5, 30.7, 122.4, and 24.5 on BLEU-4, METEOR, CIDEr, and SPICE metrics, respectively. Robustness analysis shows that PVCE maintains a high CIDEr score under structurally destructive perturbations.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.