Skip to content
Preprint

On Cross-Validation for Hyperparameter Optimization of Deep Learning Image Classifiers

Aug 2026 · 0 citations · 13 references
Computer Science

TL;DR

For small-sample medical image classification, this work recommends cross-validation-based HPO when computational resources permit because it trades additional computation for a more reliable development-time estimate of subsequent test performance.

Abstract

Hyperparameter optimization (HPO) can materially affect the performance of deep learning (DL) image classifiers, but there is little empirical guidance on how to derive the validation signal that drives it, especially for the small sample sizes common in fields such as medical imaging. We compared three HPO protocols in terms of {\em absolute performance-estimation error} (AEE; the absolute difference between the winning configuration's validation AUROC and its test AUROC): fixed holdout (F), reshuffled holdout (R), and 5-fold cross-validation (C). The search space, sampler, training procedure, architecture, and test set were held identical across protocols. We evaluated the protocols on three public datasets spanning two regimes: binary medical imaging (RSNA pneumonia radiographs and binarized HAM10000 skin lesions) and 200-class natural imaging (Tiny ImageNet), across a range of development set sizes $n$ and two backbones (ResNet-18 on all datasets, Vision Transformer (ViT-S/16) on RSNA). On the medical datasets, every point estimate favored cross-validation over both holdout protocols, with reductions in AEE largest at small sample sizes and diminishing as $n$ increased. This pattern remained robust under conservative family-wise adjustment. On Tiny ImageNet, AEE was negligible under all three protocols. Test AUROC was generally similar among protocols. Fixed holdout had lower mean AEE than reshuffled holdout in 11 of 12 medical conditions, although this secondary finding was less uniformly supported. For small-sample medical image classification, we recommend cross-validation-based HPO when computational resources permit because it trades additional computation for a more reliable development-time estimate of subsequent test performance.

View source

Similar papers

Open access Aug 2026

Cross-Architecture Assessment of Hyperparameter Optimization Techniques in Convolutional Neural Networks

It is demonstrated that hyperparameter optimization dynamics depend heavily on dataset complexity, where computational efficiency is the primary differentiator for simpler classification tasks, but optimization architecture selection becomes critical for navigating challenging medical imaging applications.

Sarab Almuhaideb, Ahmad Raza Khan · 0 citations
Review Dec 1999

Medical Image Analysis

Since the discovery of the X-ray radiation by Wilhelm Conrad Roentgen in 1895, the field of medical imaging has developed into a huge scientific discipline. The analysis of patient data acquired by current image modalities, such as computerized tomography (CT), magnetic resonance tomography (MRT), positron emission tom...

Lei Mou, Yitian Zhao, H. Fu et al. · 397 citations · ⚡29

Medical Image Analysis

UBIX can reduce their contribution to the bag-level predictions, improving reliability without retraining on new data, and potentially increases the applicability of artificial intelligence models to data from other scanners than the ones for which they were developed.

Coen de Vente, B. van Ginneken, C. Hoyng et al. · 0 citations
Open access Aug 2026

CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity

Convolutional neural networks (CNNs) are widely used for biomedical image classification, yet it remains unclear under which conditions training on reduced subsets of available data can provide reliable guidance during model development, how much training data is required to achieve stable and comparable performance ac...

G. Sgarro, Melle Mendikowski, D. Santoro et al. · 0 citations
Aug 2026

Comparative Study of CNN, Hybrid, and Transformer Architectures in Medical Image Classification.

The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.

Dina A. Elkholy, Mohamed S. Shehata, John W. Braun · 0 citations
Preprint Sep 2026

Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery

Background and Objective: External evaluation of medical-imaging AI is often collapsed into discrimination. We evaluated a computational protocol that separately tests discrimination, probability calibration, fixed operatingpoint transport, shortcut-associated signal, and limited-label recoverability for pediatric pneu...

Nazim-E-Alam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.