Skip to content

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Jul 2026 · arXiv.org · Vol abs/2607.19341 · 1 citation · 45 references
Computer Science

TL;DR

A tailored Bootstrapped Pareto Policy Optimization (BPPO) is proposed, which synergizes Bootstrapping Reward Rectification and Conflict-Aware Pareto Advantage Fusion (CPAF) and exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.

Abstract

Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shallow causal understanding, and direct knowledge recall, failing at knowledge-intensive generation. We develop \textbf{ExpertVerse}, a capability-centric benchmark to evaluate generative models via knowledge-intensive lens. ExpertVerse stratifies reasoning generation across an orthogonal taxonomy of \textit{9 cognitive capabilities} and \textit{8 expert disciplines}, yielding \textit{58 sub-disciplines}. We curate 1,611 expert-annotated instances covering single-image editing, multi-image composition, and text-to-image generation. We further develop an automated workflow to produce \textbf{ExpertVerse-100K}, a large-scale dataset with reasoning traces and knowledge-anchored rationale annotations. Based on this, we train \textbf{KnowThinker} with RL fine-tuning, a VLM reasoning engine with world knowledge that jointly generates thinking processes and refined instructions. Towards the cross-modal credit misalignment and multi-objective gradient conflicts in multi-reward optimization, we propose a tailored Bootstrapped Pareto Policy Optimization (BPPO), which synergizes Bootstrapping Reward Rectification (BRR) and Conflict-Aware Pareto Advantage Fusion (CPAF). Extensive results of both open-source and proprietary models exposes critical reasoning deficits, highlighting imperative for knowledge-intensive benchmarks towards next-generation visual generation.

View source

Similar papers

Preprint Sep 2026

Reasoning with Image Generation

Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.

Nishad Singhi, H. P. Rodriguez, Aditya Arora et al. · 0 citations
Preprint Aug 2026

REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

Chart editing requires inferring and modifying visualization code from a reference chart image based on an editing instruction, challenging fine-grained visual reasoning, instruction following, and executable code synthesis capabilities of MLLMs. Large reasoning models (LRMs) with extended Chain-of-Thought (CoT) reason...

Yuanbang Liu, Chenxi Ruan, Yihan Hou et al. · 0 citations
Preprint Aug 2026

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

This work introduces Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains, and establishes a rubric-based evaluation protocol, showing that advances in visual realism have not yet translated into reliable modeling of scientific and causal dyn...

Diandian Zhang, Tingyu Song, Linbo Fu et al. · 1 citation
Preprint Aug 2026

Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we obser...

Qiyou Liu, Yong Zhang, Jianjie Luo et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
Open access Jul 2026

Structured multi-level knowledge augmentation via small-to-large evidence-guided collaboration for knowledge-based VQA

This work proposes an inference-time evidence augmentation framework for frozen-LLM-based KB-VQA that focuses on how question-relevant multimodal evidence can be systematically constructed, refined, and organized before LLM inference.

Meng Zhang, Da-Yu Wu, Wonjun Chung · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.