Skip to content

Author

Bangrui Xu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

MoDora: A Multimodal Document AI Assistant Harness

General-purpose AI document assistants (e.g., NotebookLM) increasingly play an important role and are widely adopted across diverse domains. However, they consistently struggle on complicated multimodal documents such as financial reports and scientific papers, where hierarchical structures, complex layouts, and interleaved visual elements carry essential semantics. A key reason is that these assistants lack native multimodal data management capabilities: they typically linearize documents into flat text chunks or page images, discarding the logical hierarchy and cross-modal context that are critical for faithful evidence retrieval and reasoning. To address these limitations, we present MoDora, an interactive, tree-structured multimodal document analysis agent harness. Designed to make the document analysis process transparent and controllable, MoDora introduces an end-to-end user experience through three core functionalities: (1) an automated document ingestion engine that seamlessly transforms unstructured PDFs into a layout-aware component tree, preserving both textual hierarchy and visual elements; (2) an interactive structure visualizer that allows users to intuitively inspect the extracted document hierarchy and refine structural relations via drag-and-drop or natural language commands; and (3) a verifiable multimodal QA interface that supports complex, cross-modal queries with precise bounding-box-level grounding back to the original PDF regions, enabling users to effortlessly trace and verify the underlying evidence. Experimental results on the MMDA benchmark show that MoDora achieves an AIC-Acc of 71.1%, outperforming baselines by over 14%.

Yu-Kai Wu, Bang-Rui Xu, Shao-Lin Yu et al. · 0 citations
Jul 2026

HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research

Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existing benchmarks mainly evaluate final outcomes, such as answer correctness, report quality, or citation alignment, while providing limited visibility into whether evidence is correctly selected, linked, and aggregated into supported claims and conclusions. To address this gap, we introduce HiEviDR-Bench, a benchmark for evaluating Hierarchical Evidence Aggregation in Deep Research. HiEviDR-Bench covers open-domain and academic-domain settings under both text-only and multimodal conditions, and represents each instance with an explicit evidence graph that captures evidence selection, cross-source linking, and aggregation from evidence to intermediate claims and final conclusions. Based on this formulation, we develop a traceability-oriented evaluation framework with five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness, together with a progressive gating mechanism for fine-grained error localization. HiEviDR-Bench contains 2,000 human-validated questions with evidence graphs across multiple difficulty levels. Experiments on 16 representative multimodal large language models show that, although many systems achieve strong report quality, their performance drops markedly on citation accuracy, claim construction, and answer correctness. Further analysis shows that the main bottlenecks lie in evidence identification and intermediate claim construction, revealing that strong surface-level report quality does not necessarily imply grounded multi-stage reasoning on our benchmark.

Yubo Sun, Chunyi Peng, Yukun Yan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.