Skip to content
Open access

Learning from Prior Experiments: Meta-learning Models of Workflow Performance

Aug 2026 · SN Computer Science · Vol 7 · 0 citations · 37 references

TL;DR

This study shows that accurate performance prediction for classical machine learning workflows can be achieved through meta-learning using readily available OpenML meta-data, and indicates that careful selection of regression models is more critical than increased representational complexity.

Abstract

Evaluating the performance of machine learning workflows is a major computational bottleneck in automated machine learning (AutoML), particularly for complex pipelines involving preprocessing, model selection, and hyperparameter optimization. This work aims to develop an efficient performance prediction framework that estimates the expected accuracy of candidate machine learning workflows on unseen datasets without requiring explicit model training. We formulate performance prediction as a meta-learning regression problem that leverages historical experimental results from the OpenML platform. Machine learning workflows are represented as structured pipelines and encoded using text-based vectorization techniques, including TF-IDF, count-based, and hashing vectorizers, as well as LLM BERT embeddings. These workflow descriptors are combined with dataset-level meta-features capturing basic statistical and structural properties. Several regression models are evaluated as meta-learners, including linear models, decision trees, random forests, Gaussian processes, and gradient-boosted decision trees. The approach is systematically evaluated on the OpenML-CC18 benchmark suite using cross-validation over more than 100,000 workflow-dataset evaluations. The proposed framework achieves strong predictive performance across a wide range of workflows and datasets. In particular, gradient-boosted decision tree regressors combined with standard TF-IDF representations of workflows consistently yield the best results, reaching an average coefficient of determination \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$R^2$$\end{document} of approximately 0.8 on unseen test data. While transformer-based MiniLM embeddings were evaluated, they did not consistently outperform sparse TF-IDF representations and incurred higher computational cost. Feature ablation studies indicate that restricting vocabulary size degrades performance, while extending representations with bigrams provides only marginal gains at substantially higher computational cost. The results demonstrate robust generalization across heterogeneous workflows and dataset characteristics. This study shows that accurate performance prediction for classical machine learning workflows can be achieved through meta-learning using readily available OpenML meta-data. The proposed approach enables rapid and computationally efficient estimation of workflow performance, making it well suited for accelerating AutoML search and model selection. The results indicate that careful selection of regression models is more critical than increased representational complexity, with simple and scalable workflow encodings yielding the most robust performance. Given its scalability and flexibility, the framework provides a strong foundation for future extensions incorporating richer dataset descriptors, larger meta-datasets, and more expressive embedding and regression models.

Read PDF

Similar papers

Open access Aug 2026

In-Context Learning Meets Small Molecule Property Prediction: Benchmarking Novel Machine Learning Approaches.

Recently, a new category of machine learning approaches for tabular data has emerged: tabular foundation models (TFM), based on in-context learning. A TFM is a neural network (usually a transformer) pretrained primarily on synthetic data. Its input is an entire data set: features and labels for training records, along...

D. Matyushin, A. Sholokhova · 0 citations
Jul 2026

MLwrap: Simplifying Machine Learning Workflows in R

MLwrap is an R package that streamlines machine learning (ML) workflows, making them accessible, efficient, and reproducible, especially within the Knowledge Discovery in Databases process. It offers a unified, minimalistic interface covering all predictive modeling stages: data preprocessing, model construction, hy...

Rafael Jiménez, Javier Martínez-García, Juan José Montaño et al. · 0 citations
Open access 2024

Automated Predictive Model Selection Using Meta-Learning Techniques

The research evaluates several popular machine learning algorithms, including Decision Trees, Support Vector Machines, Random Forests, Naïve Bayes, Artificial Neural Networks, and k-Nearest Neighbor classifiers and demonstrates that meta-learning significantly improves model recommendation accuracy compared to traditio...

Pooja Agarwal, Rakesh Chandra · 0 citations
Preprint Aug 2026

Learning the Pareto Frontier of Predictive Models under Distribution Shift

Modern machine learning pipelines increasingly rely on reusing pretrained and foundation models across downstream tasks. These pretrained models can differ not only in performance but also in how they can be used: some only provide black-box predictions, while others may permit white-box access to internal representati...

Yiming Dong, Jiwei Zhao, Yang Lu · 0 citations
Review Open access Aug 2026

How to Build Machine-Learning Models for Molecular Science: A Step-by-Step, Annotated Tutorial

This tutorial provides a comprehensive, end-to-end workflow from raw data to deployed models,icitly designed for environmental chemists with limited prior experience in ML modeling while also providing practical guidance for other users seeking to strengthen their modeling workflows.

Kai Zhang, Yu-Shu Cheng, Hai-Ping Ai et al. · 0 citations
Review Open access Jul 2026

Automating Machine Learning Pipeline Design via Metalearning

This thesis introduces the Dynamic Pipeline CASH problem, which extends the CASH formulation to incorporate meta-model-driven search space creation for pipelines, using Metalearning (MtL) to dynamically build task-specific search spaces.

Edesio Alcobaça, A. Carvalho · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.