Skip to content
Preprint

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

UniHEAR is a unified lightweight framework for heterogeneous-source entity retrieval and reranking that achieves state-of-the-art retrieval and VQA performance, improving Recall@1 and Recall@1 by 6.7 and 1.2 points over the strongest baselines while maintaining a lightweight reranking architecture.

Abstract

Knowledge-Based Visual Question Answering (KB-VQA) requires retrieving entity knowledge from external sources to answer visually grounded questions. Existing retrieval-augmented systems suffer from two critical limitations. First, relying on a single retrieval modality creates a Single-Source Retrieval Bottleneck, missing ground-truth entities that are only accessible through complementary sources. Second, dual-tower pointwise rerankers suffer from Retrieval-Source-Blind Reranking, as they overlook retrieval origins and candidate-level retrieval priors, leading to redundant modality reliance. To address these challenges, we propose UniHEAR, a unified lightweight framework for heterogeneous-source entity retrieval and reranking. UniHEAR constructs a Coarse Retrieval Descriptor for each candidate entity, and introduces Retrieval-Guided Attentive Modality Gating to condition modality attention weights on this descriptor, complemented by Entropy-Weighted Source Fusion of coarse retrieval priors. A hybrid training strategy combining contrastive learning with an auxiliary modality-preserving loss unifies entity-level and section-level retrieval within a single model. Extensive experiments on E-VQA and InfoSeek demonstrate that UniHEAR achieves state-of-the-art retrieval and VQA performance, improving Recall@1 by 6.7 and 1.2 points over the strongest baselines while maintaining a lightweight reranking architecture. Code and model are available at https://github.com/iven-luo/UniHEAR.

View source

Similar papers

Preprint Aug 2026

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

KBMR is proposed, the first MLLM-based embedding retriever tailored for KB-VQA, and an MLLM-based semantic discriminator that generates continuous entity-consistency weights is introduced to tackle the challenge of noisy supervision in Wikipedia-scale retrieval.

Hangrui Xu, Zheng-Xian Wu, Yu Yu et al. · 0 citations
Conference Open access Sep 2026

Dynamic Multi-Path Retrieval for Knowledge-based Visual Question Answering

Dynamic Multi-Path Retrieval for KB-VQA (DMRAG) is proposed, which re-trieves candidates through multiple retrieval paths that capture complementary visual and semantic cues and performs Question-Adaptive Gated Fusion to balance contributions from different modalities according to the query’s information need.

Zeyu Song, Yimin Deng, Yu-Xin Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval

This work proposes a training-free multi-hop retrieval framework that integrates evidence-conditioned exploration, passage-specific contrastive refinement, and coverage-aware final ranking and demonstrates consistent improvements in retrieval quality and downstream QA performance over baselines.

Jungmin Yun, Youngbin Kim · 0 citations
Aug 2026

A visual question answering model based on entity knowledge selection

EKS is a novel framework that leverages entity relations in commonsense knowledge graphs to dynamically generate knowledge sentences relevant to both visual and textual entities and formulates knowledge selection as a relevance scoring problem, where semantic similarity is used to measure the relevance between knowledg...

Kun Zhu, Kun Zhou, De-Xin Zhao · 0 citations
Preprint Aug 2026

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

This work introduces UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT), and aligns the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning.

Ze-Long Sun, Jun Wang, Kaicheng Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval

UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence and achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.

Libiao Chen, Xiyang Liu, Yanheng Wei et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.