Adaptive versus non-adaptive sampling for Gaussian-RBF surrogates: a replicated benchmark across analytical and differential-equation models
Adaptive sampling extends a nested design point by point, but repeated surrogate fitting and acquisition search are justified only when they improve prediction or reduce expensive model evaluations. This paper presents a controlled, replicated comparison of nine sampling strategies within a fixed Gaussian radial-basis-function (RBF) pipeline and examines when sequential acquisition is justified. Nine strategies—Random, Latin-hypercube sampling (LHS), scrambled Halton, sequential maximin, P-greedy, a nearest-neighbour leave-one-out (NN-LOO) proxy, a β-NN-LOO/P proxy, and RBF criteria in the spirit of MEPE and EIGF—are compared on seven analytical and differential-equation problem/QoI combinations: the smooth analytic Branin function; a steady heat problem with a discontinuous conductivity (Heat); an advection-dominated Graetz problem with a boundary layer (Graetz); and a two-species reaction–diffusion system with two (2S-RD-2P) or four (2S-RD-4P) active parameters, each with a reaction-ODE functional ψ1 and a reaction–diffusion state functional ψ2. Thirty design replications use a common 2000-point validation design. Supported endpoint improvements in the original adaptive-versus-static comparison require agreement between Holm-adjusted permutation and paired Wilcoxon analyses; tolerance performance combines attainment probability with the conditional first evaluated budget. No strategy dominates within the tested suite and budgets. Geometry-retaining adaptive criteria improve Heat, Graetz and both four-parameter 2S-RD QoIs, whereas LHS remains effective on both two-parameter QoIs. Sequential maximin is best by endpoint median on Branin and 2S-RD-4P/ψ2, second on 2S-RD-4P/ψ1, but eighth on Graetz, so response-informed sampling does not uniformly dominate strong nested geometry. Heat accuracy changes materially over εscore ∈ {0.5, 1, 2}, and fitting the physical rather than logarithmic Heat target worsens every compared method’s physicalscale endpoint median. At Nc = 4000 the four-dimensional candidate pool is much coarser than its two-dimensional counterpart; quadrupling the 2S-RD-4P pool gives small, non-systematic endpoint shifts but less stable detailed rankings. The resulting decision map is a scoped benchmark-based guide; its sensitivity to the scoring parameter, fitted-target scale and finite candidate pool is stated explicitly.