Skip to content
Review Open access

Evaluating Machine Learning Models in Nonstandard Settings: An Overview and New Findings

Aug 2026 · Statistical Science · Vol 41, pp. 524-560 · 0 citations · 55 references

TL;DR

The findings corroborate the concern that standard resampling methods often yield biased GE estimates in nonstandard settings, underscoring the importance of tailored GE estimation.

Abstract

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in nonstandard settings, particularly those where observations are not independently and identically distributed, resampling using simple random data divisions may lead to biased GE estimates. This paper strives to present well-grounded guidelines for GE estimation in various such nonstandard settings: clustered data, spatial data, unequal sampling probabilities, concept drift and hierarchically structured outcomes. Our overview combines well-established methodologies with other existing methods that, to our knowledge, have not been frequently considered in these particular settings. A unifying principle among these techniques is that the test data used in each iteration of the resampling procedure should reflect the new observations to which the model will be applied, while the training data should be representative of the entire data set used to obtain the final model. Beyond providing an overview, we address literature gaps by conducting simulation studies and a real data study. These studies assess the necessity of using GE-estimation methods tailored to the respective setting. Our findings corroborate the concern that standard resampling methods often yield biased GE estimates in nonstandard settings, underscoring the importance of tailored GE estimation.

Read PDF

Similar papers

Preprint Aug 2026

NOMADD: Numerical Optimization of Models Adapting to Data Drift

This paper offers an alternative post-hoc method to reduce concept drift, which is applicable to a variety of models, from trees to neural networks to tabular foundation models, and explores the promise and challenges of extending this tool to other modalities.

Swapn Shah, K. Burghardt · 0 citations
#machine learning Review Sep 2026

Machine Learning under Imperfect Data: Challenges and Methods

Machine-learning models are commonly developed under an assumption that training and test data are sufficiently complete, balanced, labelled, and drawn from compatible distributions. In practice, one or more of these conditions is often violated. Measurements may be missing or corrupted, rare classes may be poorly represented, supervision may be weak, and the deployment environment may differ from the training environment. These imperfections are usually treated as separate technical problems, although they alter learning through a small number of shared mechanisms: loss of information, biased empirical risk, ambiguous supervision, and unstable representations. This short survey organises representative methods around these mechanisms. It reviews reconstruction and generation, rebalancing and representation calibration, learning with limited supervision, adaptation across domains and modalities, and reliability under distribution change. The discussion highlights the limits of plausible reconstruction, benchmark-specific correction, and adaptation without trustworthy feedback. It concludes with directions for evidence-aware learning, uncertainty-preserving prediction, and evaluation that separates visual plausibility from decision utility.

Masoumeh Zareapoor · 0 citations
Nov 2026

Theoretical analysis of transfer learning for temporally dependent observations

Transfer learning has become increasingly important in recent years as it enables models to quickly adapt to new tasks, environments and data sets. It can improve the accuracy of machine learning models and reduce the training time required for them. However, theoretical foundation of transfer learning algorithms is rather scarce, particularly for time series models. In this work, we focus on high-dimensional vector autoregressive models and provide a two-step procedure to conduct transfer learning utilizing auxiliary data sets followed by constructing confidence intervals for model parameters in the high-dimensional regime. Further, a new method to select informative sets from auxiliary sets is introduced. Finally, two debiasing techniques are developed to perform inference for model parameters. Theoretical properties of all proposed algorithms are established under mild conditions which allow for heavy-tailed distributions and dependence among auxiliary and target data sets. Given certain level of similarity between the informative models and the target model, it is shown that the proposed algorithm achieves the minimax rate. Lastly, the empirical performance of proposed methods is tested through analyzing both simulated data as well as an EEG data set.

Abolfazl Safikhani, Mingliang Ma · 0 citations
Open access 2026

Novel Regularization Methods to Prevent Overfitting in Machine Learning Models

The problem of overfitting is one of the most persistent in the current machine learning (ML), especially as models and data dimensionality increases. Although classical regularization methods like L1, L2, dropout, and early stopping have been proven to be efficient, their weakness is realized in large-scale, deep, and data-sparse learning settings. In the current paper, a detailed study has been carried out on new regularization techniques that aim to enhance the performance of generalization and, at the same time, ensure the expressiveness of the model. We present a single taxonomy of the new regularization methods such as adaptive regularization, information-theoretic constraints, structured sparsity, stochastic regularization and regularization at the representation level. Moreover, we suggest Hybrid Adaptive Information Regularization (HAIR) which is a dynamic complex/generalization balance which is regularized by entropy-based penalties and parameter-sensitivity analysis. Numerous comparative studies show that the suggested approach is more effective than the traditional approaches in various learning paradigms. The findings have emphasized the importance of advanced regularization in developing robust, scalable and interpretable ML systems. The current study provides a certain contribution to both theoretical background and methodological developments as well as empirical findings in favor of next-generation regularization approaches.

Rak esh, A. An · 0 citations
#artificial intelligence Preprint Sep 2026

Learning with Synthetic Data via SGD in High-Dimensional Linear Regression

Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper, we study how synthetic data affects the generalization of one-pass SGD in high-dimensional linear regression with model shift. We establish finite-sample risk bounds for mixed and two-stage training, separating standard bias and variance from source-mismatch effects, namely fluctuation and persistent drift under mixing and filtered initialization bias under two-stage. These bounds reveal a sharp contrast: mixed training induces strong model collapse, while two-stage training avoids the floor by using synthetic data only in the first stage, showing that collapse is not inevitable under a simple data curriculum. Under a random sketch model, we further obtain scaling laws for both protocols, with tight results for mixed training in the optimization-saturated regime. These laws show that larger models may amplify synthetic-induced degradation under mixing, and quantify how high-quality synthetic pretraining may reduce bias in two-stage training. Finally, we establish an exact finite-sample necessary-and-sufficient condition for two-stage training to strictly outperform real-only training under the same real-data budget and identical real-stage updates. Overall, our results highlight that synthetic data is neither inherently harmful nor beneficial; its effect depends critically on both its quality and the training protocol used to incorporate it.

Jichu Li, Di-Fan Zou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.