Skip to content
Open access

Dual Text-guided Cross-attention for Global–local Visual Fusion in Vietnamese Visual Question Answering

Aug 2026 · Asian Journal of Research in Computer Science · Vol 19, pp. 1-11 · 0 citations

TL;DR

This study proposes a dual-stream architecture that exploits complementary Transformer-based and convolutional visual representations for Vietnamese VQA, and shows improvements over the corresponding single-stream convolutional baselines, supporting the complementary role of the two visual representations.

Abstract

Visual Question Answering (VQA) requires effective cross-modal reasoning between visual content and natural-language questions. This challenge is particularly significant for Vietnamese due to the relatively limited availability of annotated VQA resources. This study proposes a dual-stream architecture that exploits complementary Transformer-based and convolutional visual representations for Vietnamese VQA. A Vision Transformer (ViT) is employed to obtain globally contextualised visual features, while ConvNeXt V2 preserves spatially structured visual information. Questions are encoded using PhoBERT. Instead of directly merging the two visual streams, the proposed model uses the textual representation as a shared query in two independent cross-attention modules, allowing question-relevant information to be retrieved separately from each visual representation prior to fusion. The resulting representations are concatenated along the sequence dimension, pooled, and combined with the sentence-level textual embedding through a residual connection for answer classification. Experiments on the ViVQA benchmark show that the proposed architecture achieves 64.31% accuracy and 62.65% F1-score. Additional experiments across multiple ConvNeXt V2 backbone scales consistently show improvements over the corresponding single-stream convolutional baselines, supporting the complementary role of the two visual representations.

Read PDF

Similar papers

Open access Aug 2026

Indonesia–English Bilingual Visual Question Answering Using Partial Fine-Tuning on a ViT-GPT2 Architecture

Visual Question Answering (VQA) is a multimodal task that integrates visual understanding and natural language processing to generate answers based on information contained in an image. Most existing VQA research focuses on the English language and general-domain datasets, limiting its applicability to bilingual enviro...

A. Anas, Hazriani Hazriani, Yuyun Yuyun et al. · 0 citations
Aug 2026

Cross-modal alignment enhancement for lightweight large vision language models

A Low-Complexity Cross-Modal Alignment via Projection (LCAP) network is proposed, which introduces Projective Token Compression (PTC), which leverages Mish activation and adaptive average pooling to reduce feature redundancy while enhancing discriminative information, and Positional Spatial Enhancement (PSE), which exp...

Yu-Chen Sha, Lingli Wan, Ge Yang et al. · 0 citations
#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
Open access 2026

Decoupled global-local collaborative network for visual question answering

Visual Question Answering (VQA) aims to achieve cross-modal semantic understanding through joint modeling of visual content and natural language. Although existing attention-based approaches effectively align features, they struggle to simultaneously accommodate global semantic modeling and local fine-grained perceptio...

Gan-Long Zhou, Dezhi Han, Xiang Shen et al. · 0 citations
Preprint Sep 2026

ConvCue: Complementary Visual Inductive Biases for Vision-Language Models

Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations can improve existin...

Zi-Xuan Lan, Shi-Chu Sun · 0 citations
Conference Aug 2026

ViG: Grounded Adaptation of Transformer-Based Models for Low-Resource Vietnamese Image Captioning

Dual-feature captioners that fuse region-level and grid-level visual features define the state of the art in image captioning, yet they owe this strength to visual encoders trained on massive English corpora. A low-resource language such as Vietnamese cannot retrain these encoders, and finetuning the decoder alone cann...

Hieu Thien Vu, N. H. Ngoc, Thang Cap Pham Dinh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.