Skip to content
#graph neural networks Review Open access

Exploring multi-answer visual question answering with object detection: a systematic review

Sep 2026 · International Journal of Informatics and Communication Technology (IJ-ICT) · Vol 15, pp. 1097 · 0 citations · 82 references

TL;DR

This systematic literature review (SLR) focuses on multi answer VQA systems and the use of object detection, following the PRISMA 2020 guidelines and proposes a taxonomy of multi-answer VQA organized along four dimensions.

Abstract

Visual question answering (VQA) is a challenging research area that enables machines to answer natural language questions based on visual content by jointly understanding images and text. Conventional VQA systems typically produce a single answer for each image–question pair. However, many real world visual questions are ambiguous or complex, allowing multiple valid answers to exist. This systematic literature review (SLR) focuses on multi answer VQA systems and the use of object detection, following the PRISMA 2020 guidelines. We analyzed 58 peer-reviewed journal articles retrieved from the Scopus database published between 2020 and 2025. Ten of these studies clearly stated that generating multiple answers was their main goal. Forty-eight others indirectly supported answer variability by using object-based or multi-instance reasoning. Through this review, we examine the current methodologies for supporting multi-answer generation, including model architecture, datasets, and evaluation metrics. Most multi answer generation approaches utilize attention mechanisms, graph neural networks, and transformer-based models. Additionally, we propose a taxonomy of multi-answer VQA organized along four dimensions. Limitations are identified in datasets and evaluation metrics (i.e., answer ambiguity/subjectivity). Future research should focus on improving model interpretability and designing an evaluation framework that incorporates subjective and context-sensitive responses.

Read PDF

Similar papers

Conference Open access Sep 2026

Dynamic Multi-Path Retrieval for Knowledge-based Visual Question Answering

Dynamic Multi-Path Retrieval for KB-VQA (DMRAG) is proposed, which re-trieves candidates through multiple retrieval paths that capture complementary visual and semantic cues and performs Question-Adaptive Gated Fusion to balance contributions from different modalities according to the query’s information need.

Zeyu Song, Yimin Deng, Yu-Xin Zhang et al. · 0 citations
Book Open access Aug 2026

SciChart: Visual Question Answering and Reasoning for Scientific Spectral Chart

Charts play a key role in scientific research, offering a concise and visual way to present complex data. For Multimodal Large Language Models (MLLMs), the ability to comprehend charts is critical, as it requires both visual perception and reasoning that bridges graphical and textual information. However, existing char...

Tan Yue, Rui Mao, Xuzhao Shi et al. · 1 citation
Aug 2026

A visual question answering model based on entity knowledge selection

EKS is a novel framework that leverages entity relations in commonsense knowledge graphs to dynamically generate knowledge sentences relevant to both visual and textual entities and formulates knowledge selection as a relevance scoring problem, where semantic similarity is used to measure the relevance between knowledg...

Kun Zhu, Kun Zhou, De-Xin Zhao · 0 citations
Preprint Aug 2026

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we obser...

Qiyou Liu, Yong Zhang, Jianjie Luo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmarks often target predefined capabilities, such as multi-hop retrieval or long-form synthesis, whereas users ask photo-grounded questions span...

Hao-Nan Jiang, Guo-Jian Zhan, Jian-Cong Xie et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.