General-purpose AI document assistants (e.g., NotebookLM) increasingly play an important role and are widely adopted across diverse domains. However, they consistently struggle on complicated multimodal documents such as financial reports and scientific papers, where hierarchical structures, complex layouts, and interleaved visual elements carry essential semantics. A key reason is that these assistants lack native multimodal data management capabilities: they typically linearize documents into flat text chunks or page images, discarding the logical hierarchy and cross-modal context that are critical for faithful evidence retrieval and reasoning. To address these limitations, we present MoDora, an interactive, tree-structured multimodal document analysis agent harness. Designed to make the document analysis process transparent and controllable, MoDora introduces an end-to-end user experience through three core functionalities: (1) an automated document ingestion engine that seamlessly transforms unstructured PDFs into a layout-aware component tree, preserving both textual hierarchy and visual elements; (2) an interactive structure visualizer that allows users to intuitively inspect the extracted document hierarchy and refine structural relations via drag-and-drop or natural language commands; and (3) a verifiable multimodal QA interface that supports complex, cross-modal queries with precise bounding-box-level grounding back to the original PDF regions, enabling users to effortlessly trace and verify the underlying evidence. Experimental results on the MMDA benchmark show that MoDora achieves an AIC-Acc of 71.1%, outperforming baselines by over 14%.
Yu-Kai Wu, Bang-Rui Xu, Shao-Lin Yu et al.· Proceedings of the VLDB Endo...· 0 citations
OmniOpt supplies the research community with an operational coordinate system for selecting optimizers under explicit mechanism and objective assumptions, and charts a direction for the future development of the optimizer community.
Siyuan Li, Jiabao Pan, Yumou Liu et al.· 0 citations
Results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.
Boxiu Li, Zimo Wen, Yijia Fan et al.· 2 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.