Skip to content
Book Open access

BioWeaver: Adaptive Workflow Orchestration for Biomedical Data Integration with Progressive Deep Web Exploration

Aug 2026 · International Conference on Statistical and Scientific Database Management · pp. 1-12 · 0 citations · 19 references
Computer Science

TL;DR

BioWeaver is presented, an adaptive workflow orchestration system that converts natural-language biomedical questions into executable scientific workflow graphs and suggests that adaptive workflow graphs provide a practical foundation for reproducible, source-grounded biomedical data integration across heterogeneous and deep web resources.

Abstract

Biomedical researchers often need to combine evidence from heterogeneous sources, including REST APIs, form-based deep web databases, semi-structured web pages, and downloadable files. Existing workflow systems support reproducible analysis but require manually specified pipelines, while LLM-based agents can issue tool calls but often lack explicit data dependencies, provenance, and controlled exploration. We present BioWeaver, an adaptive workflow orchestration system that converts natural-language biomedical questions into executable scientific workflow graphs. Each graph represents typed retrieval, refinement, and synthesis steps over a catalog of heterogeneous connectors. During execution, BioWeaver uses Progressive Data Refinement (PDR) to expand the workflow when intermediate results reveal useful follow-up records, while a composite information-gain heuristic and human-in-the-loop checkpoints control exploration depth. A unified connector abstraction supports APIs, browser-automated forms, HTML tables, and downloadable files under a common execution model. We evaluate BioWeaver on 45 queries across two biomedical question sets with gold-standard answers. Results show that BioWeaver improves overall answer quality by 6.2–14.5% over LLM baselines and by 41.4–53.8% over a ReAct agent. Ablation studies show that runtime refinement, planner knowledge, and information-gain stopping each contribute to system effectiveness. These results suggest that adaptive workflow graphs provide a practical foundation for reproducible, source-grounded biomedical data integration across heterogeneous and deep web resources.

Read PDF

Similar papers

Review Open access Sep 2026

Accelerating metadata annotation in collaborative research centers: A hybrid AI workflow for biomedical entities

Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verifica...

M. Watter, F. Engel, Aref Kalantari et al. · 0 citations
Preprint Oct 2026

Bridging the Omics Divide: A Modular Relational Approach to Multi-Layer Biological Data Management

Background: Rapid growth of high-throughput molecular data demands systems for efficient retrieval, integration, and scalability across omics layers. Traditional file-based workflows hinder cross-modal analysis and reproducibility because of fragmented storage and ad-hoc querying. Few existing tools for genomic variati...

Alessandro Balestrucci, Donald Friggieri, Andrea Gariboldi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OpenAl4S: Code as Action, Science as Sessions

The results suggest that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows, and recommend that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows.

Gong-Bo Zhang, Hao Li, Yu Wang et al. · 0 citations
Review

Validation-Driven Automation in a Federated XML Ingest Pipeline: The NCBI Bookshelf Case

This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.

Kin Ng, Lisandro Gonzalez, Stacy Lathrop · 0 citations
#artificial intelligence Preprint Sep 2026

Bioinfoysis Technical Report

A multi-agent harness that represents each request as a persistent, artifact-grounded analysis run, and demonstrates that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow.

Qi Shao, Xin Zhang, Zhou-Yang Yuan et al. · 0 citations
Review Open access Aug 2026

SilkRoute: A Descriptor-Driven Framework for Reproducible Multi-Source Biomolecular Data Acquisition

Background Biomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedure...

Diego Fernández, Julián García-Vinuesa, Diego Alvarez-Saravia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.