Skip to content
Preprint

CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications

Aug 2026 · 0 citations · 49 references
Computer Science Physics

TL;DR

The system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets are described.

Abstract

CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and technological applications. CheMLFlow targets a common bottleneck in scientific machine learning development, where researchers often need to assemble data acquisition, curation, representation, model training, validation, screening, interpretation, and reporting into a reproducible pipeline, even when their primary research contribution concerns only one stage. CheMLFlow provides modular workflow components, ready-to-run reference pipelines, standardized artifacts, and evaluation outputs that reduce orchestration overhead and support benchmarking across methods and datasets. The platform is designed to be extensible, reproducible, and automation friendly, with pluggable representations and models, deterministic splits, explicit run artifacts, batch execution, and report generation. As scientific software increasingly moves toward agent assisted experimentation, CheMLFlow's configuration driven workflows and structured outputs also provide a practical interface for coding agents to help users construct experiments, inspect results, and summarize findings under human supervision. This article describes the system architecture, core workflows, and benchmarks that reach literature performance for quantum mechanical, physicochemical and bioactivity property prediction, and use cases involving time series datasets demonstrating applications beyond molecular chemistry datasets.

View source

Similar papers

Open access Sep 2026

PromptBio: An Agentic Platform for End-to-End Computational Biomedical Research

Modern biomedical research increasingly depends on complex computational analyses, yet translating a scientific question into a reliable workflow still requires substantial technical expertise and manual coordination. PromptBio is a multi-agent AI platform available through a web portal at https://promptbio.ai that add...

Min-Zhe Zhang, Wen-Hao Gu, Bo-Wei Han et al. · 0 citations
Review Open access Aug 2026

TaxoFlow: a step-by-step tutorial to build a nextflow pipeline for metagenomics taxonomic classification

An open, interactive and web-based tutorial that guides scholars with basic command-line skills through the detailed development of a validated and reproducible Nextflow metagenomics classification pipeline, which aims to lower technical barriers in microbiome bioinformatics and promote best practices in metagenomics d...

Jeferyd Yepes-García, Laurent Falquet · 0 citations
Open access Aug 2026

ChemLint: Conversational Cheminformatics with Large Language Models

We introduce ChemLint, an open-source Model Context Protocol (MCP) server that connects any MCP-compatible large language model to a curated suite of local cheminformatics and machine learning tools, enabling rigorous molecular data handling through a conversational interface. Molecular machine learning studies are f...

D. van Tilborg, Francesca Grisoni · 0 citations
Book Open access Aug 2026

Benchmarking LLM Agents on Real-World Biological Database Curation for Data-Driven Scientific Discovery

BioDataLab evaluates the capability of autonomous agents to transform raw, heterogeneous biological resources into structured, analysis-ready databases, and underscores that while LLMs are proficient in downstream reasoning, autonomous upstream curation remains a formidable frontier.

Jiaxian Yan, Xi Fang, Jintao Zhu et al. · 0 citations
Preprint Aug 2026

ChemReporter: A Framework for Curating and Exporting Large-Scale Chemical Datasets for MLIP Training

Training set quality and diversity are key determinants of the reliability of machine learning interatomic potentials (MLIPs), yet using massive datasets in full is often impractical and redundant, making intelligent data selection essential. A major bottleneck, however, is the lack of infrastructure for uniformly acce...

M. Bluntzer, Jules Tilly, Christoph Brunken · 0 citations
Preprint Aug 2026

ALKEMIE Agent: an autonomous platform for computational materials design

ALKEMIE Agent is introduced, an agentic platform in which retrieval-augmented generation, a materials-computation knowledge base, registered skills, database-supported provenance, AI-assisted structure modeling, bounded task execution, tool-calling iteration, and error-diagnostic assistance are integrated within a trac...

Hongfu Huang, Yu-Zhe Li, Ao Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.