Aug 2026· International Conference on Statistical and Scientific Database Management· pp. 1-12· 0 citations· 19 references
Computer Science
TL;DR
BioWeaver is presented, an adaptive workflow orchestration system that converts natural-language biomedical questions into executable scientific workflow graphs and suggests that adaptive workflow graphs provide a practical foundation for reproducible, source-grounded biomedical data integration across heterogeneous and deep web resources.
Abstract
Biomedical researchers often need to combine evidence from heterogeneous sources, including REST APIs, form-based deep web databases, semi-structured web pages, and downloadable files. Existing workflow systems support reproducible analysis but require manually specified pipelines, while LLM-based agents can issue tool calls but often lack explicit data dependencies, provenance, and controlled exploration. We present BioWeaver, an adaptive workflow orchestration system that converts natural-language biomedical questions into executable scientific workflow graphs. Each graph represents typed retrieval, refinement, and synthesis steps over a catalog of heterogeneous connectors. During execution, BioWeaver uses Progressive Data Refinement (PDR) to expand the workflow when intermediate results reveal useful follow-up records, while a composite information-gain heuristic and human-in-the-loop checkpoints control exploration depth. A unified connector abstraction supports APIs, browser-automated forms, HTML tables, and downloadable files under a common execution model. We evaluate BioWeaver on 45 queries across two biomedical question sets with gold-standard answers. Results show that BioWeaver improves overall answer quality by 6.2–14.5% over LLM baselines and by 41.4–53.8% over a ReAct agent. Ablation studies show that runtime refinement, planner knowledge, and information-gain stopping each contribute to system effectiveness. These results suggest that adaptive workflow graphs provide a practical foundation for reproducible, source-grounded biomedical data integration across heterogeneous and deep web resources.
Collaborative Research Centers rely on FAIR-compliant, richly structured metadata, yet manual annotation is a major bottleneck. We implemented a search-augmented large language model (LLM) workflow within a local research data management system to pre-annotate biomedical entities, using human-in-the-loop verifica...
M. Watter, F. Engel, Aref Kalantari et al.· Bioinformatics Advances· 0 citations
Background: Rapid growth of high-throughput molecular data demands systems for efficient retrieval, integration, and scalability across omics layers. Traditional file-based workflows hinder cross-modal analysis and reproducibility because of fragmented storage and ad-hoc querying. Few existing tools for genomic variati...
Alessandro Balestrucci, Donald Friggieri, Andrea Gariboldi et al.· 0 citations
The results suggest that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows, and recommend that integrating persistent execution with session-level provenance can improve the reliability of AI-assisted scientific workflows.
Gong-Bo Zhang, Hao Li, Yu Wang et al.· 0 citations
This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.
Kin Ng, Lisandro Gonzalez, Stacy Lathrop· Balisage Series on Markup Te...· 0 citations
A multi-agent harness that represents each request as a persistent, artifact-grounded analysis run, and demonstrates that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow.
Qi Shao, Xin Zhang, Zhou-Yang Yuan et al.· 0 citations
Background Biomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedure...
Diego Fernández, Julián García-Vinuesa, Diego Alvarez-Saravia et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.