Dual Text-guided Cross-attention for Global–local Visual Fusion in Vietnamese Visual Question Answering
This study proposes a dual-stream architecture that exploits complementary Transformer-based and convolutional visual representations for Vietnamese VQA, and shows improvements over the corresponding single-stream convolutional baselines, supporting the complementary role of the two visual representations.