Skip to content
Open access

Intelligent Prompt Construction for Large Language Models in Knowledge-based Visual Question Answering

Jul 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · Vol 22, pp. 1-21 · 0 citations · 52 references

TL;DR

The Intelligent Prompt Construction Framework (IPCF), which equips an autonomous agent with the ability to dynamically generate task-specific prompts, and achieves performance gains over existing baselines on the OK-VQA and A-OKVQA datasets.

Abstract

Large Language Models (LLMs) have demonstrated strong capabilities in knowledge-based Visual Question Answering (VQA). However, existing prompt construction methods are often rigid and fail to fully exploit the reasoning potential of LLMs. To address this limitation, we propose the Intelligent Prompt Construction Framework (IPCF), which equips an autonomous agent with the ability to dynamically generate task-specific prompts. IPCF consists of a planner and a toolbox: the planner, powered by an LLM, enables autonomous decision-making, while the toolbox provides three tools—the vanilla VQA model for inspiration, the LLM for knowledge injection, and a knowledge base for information retrieval. This architecture allows the agent to flexibly determine when and how to invoke each tool and to construct adaptive prompts accordingly. Experimental results show that IPCF achieves performance gains of 2.6 and 1.9 points over existing baselines on the OK-VQA and A-OKVQA datasets, respectively.

Read PDF

Similar papers

Generating then Refining for Reliable Knowledge Base Question Answering

Evaluations on standard KBQA benchmarks show that the proposed ARI-KBQA enhances model performance with a reduced search space, especially in complex multi-hop query scenarios.

Jian-Qi Gao, Hang Yu, Jian Cao et al. · 0 citations
Open access Jul 2026

Structured multi-level knowledge augmentation via small-to-large evidence-guided collaboration for knowledge-based VQA

This work proposes an inference-time evidence augmentation framework for frozen-LLM-based KB-VQA that focuses on how question-relevant multimodal evidence can be systematically constructed, refined, and organized before LLM inference.

Meng Zhang, Da-Yu Wu, Wonjun Chung · 0 citations
Aug 2026

A visual question answering model based on entity knowledge selection

EKS is a novel framework that leverages entity relations in commonsense knowledge graphs to dynamically generate knowledge sentences relevant to both visual and textual entities and formulates knowledge selection as a relevance scoring problem, where semantic similarity is used to measure the relevance between knowledg...

Kun Zhu, Kun Zhou, De-Xin Zhao · 0 citations
Preprint Aug 2026

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we obser...

Qiyou Liu, Yong Zhang, Jianjie Luo et al. · 0 citations
Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Preprint Aug 2026

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

This work proposes SAFE-G, a Structure-Aware Faithful Evidence-guided Generation framework, which enables precise evidence localization and trustworthy reasoning and introduces a reinforcement learning strategy with an evidence-grounded reward that assigns credit to correct answers only when the selected evidence is co...

Long Shu, Shuochen Liu, Wei Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.