A novel spatial-frequency invariant semantic learning model designed to overcome limitations in visual question answering models by incorporating frequency-domain features as an additional modality and employing invariant feature learning techniques, effectively reduces bias without relying on external datasets.
Abstract
Visual question answering models often encounter challenges related to data biases and exhibit limited performance in specialized domains such as cultural heritage, where recognizing fine-grained textures is crucial. In this study, we propose a novel spatial-frequency invariant semantic learning model designed to overcome these limitations. By incorporating frequency-domain features as an additional modality and employing invariant feature learning techniques, our model effectively reduces bias without relying on external datasets. The proposed model extracts invariant representations across textual, spatial, and frequency domains, thereby filtering out spurious correlations. Comprehensive experiments on benchmark datasets, including VQA-CP v2 and GQA-OOD, demonstrate that our model achieves state-of-the-art results. Furthermore, our model exhibits enhanced robustness when applied to cultural heritage datasets, proficiently handling complex visual textures and multimodal reasoning tasks. This model enhances the capabilities of visual question answering systems in identifying artistic materials and techniques, providing a robust solution tailored to domain-specific applications.
ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.
A novel training framework that enhances counterfactual contrastive learning for VQA and introduces a three-stage curriculum for stable multi-objective optimization, an enhanced Batch-Contrastive loss for more discriminative feature learning and two novel regularizers.
Truong-Binh Duong, T. Tran, Ngoc-Thao Nguyen et al.· 0 citations
Visual Question Answering (VQA) is a challenging cross-disciplinary task that combine natural language processing and computer vision together. VQA answer natural language questions based on images. Current VQA systems prone to generate plausible but incorrect responses without indicating uncertainty. To address this l...
Prakhar Shukla, Ankit Kumar, Pulkit Singh et al.· 2026 4th International Confe...· 0 citations
A dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation is introduced.
For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.
Jiaang Li, Chengzu Li, Zhaochong An et al.· arXiv.org· 0 citations
KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.
Hangrui Xu, Zheng-Xian Wu, Yu Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.