Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample
We build an instrument that reads, from a single fit and with no oracle, whether the operator a hybrid PDE-parameter estimator postulates is wrong-and separates that from a merely unidentifiable parameter. On one self-adjoint parabolic inverse problem, an information-matrix statistic with plug-in scale and per-seed parameter has median 0.19 under correct specification, rejection rate $0.033$ against a pre-registered ceiling of $0.10$, and rises to $224$ and $85$ under two misspecifications, firing in every replicate. On a correctly specified but non-identifiable design it stays mute-$0.050$ at $n=200$, Clopper-Pearson $[0.024, 0.090]$-while a rank statistic collapses to zero at a pre-registered boundary $c_5^*=2.15\times10^{-3}.$ Two readings of one fit therefore separate the two failures across the three designs a deployable test reaches. That separation is the contribution; detection alone is a crowded flank. In sample it is a bound, out of sample a direction. It is needed because the usual accuracy check is blind: the misspecified estimator's in-domain RMSE is $2.7\times 10^{-2}$, below the observation noise for $\sigma\geq 0.05,$ while the coefficient is wrong by $29.7\%$ at zero noise, $31.2\%$ at the loudest. Nor is the failure architectural: a one-parameter curve fit, a bare parameter and multilayer perceptrons of $49$ and $241$ parameters converge to the same pseudo-true, matched in closed form to $0.07\%,$ whereas a physics-informed network, with its composite objective, converges to a disjoint one. We report where the instrument is blind, a pre-registered negative where a neural estimator loses to Tikhonov-regularized inversion at recovery, and the hypothesis under which its guarantee holds but a trained network violates it.
Instrumental-variables estimation increasingly pools many or high-dimensional instruments into a single machine-learned first stage, with rich controls partialled out. The resulting estimand, the partialled-out IV coefficient built from any signal of the instruments, is a signal-weighted average of the heterogeneous effects, which gives an opaque first stage a precise structural meaning. The average is convex whenever a covariance-monotonicity condition holds, and we provide a microfoundation for that condition based on vector monotonicity. With a learned signal, however, the usual debiased moment is not Neyman-orthogonal, and its first-order bias is a drift toward the learner's own signal-weighted average, so naive inference remains valid only for that learner-dependent target. We construct a heterogeneity-robust orthogonal score that restores $\sqrt{N}$ inference on the fixed, learner-invariant target at no efficiency cost, and provide a Hausman-type diagnostic and identification-robust confidence sets.
Accurate option prices do not imply accurate recovery of the latent risk-neutral density. We study this distinction with two complementary benchmarks. A controlled benchmark exposes simulator-truth densities for latent evaluation, while a chronological NIFTY benchmark tests only held-out market prices. A two-component lognormal mixture has the lowest aggregate price, $L^1$, Wasserstein, and fixed-tail errors on the synthetic benchmark. Learned operators retain narrower strengths: DeepONet reduces 1% quantile and variance error by 39.0% and 34.6% relative to the mixture, and a quote transformer reduces $L^1$ by 16.4% on the structurally misspecified Merton family. A numerical conditioning analysis explains why these rankings can differ: after enforcing mass and forward constraints, 95 of 126 pricing directions are numerically null, and two densities separated by $L^1 = 0.061$ produce identical prices on the covered strikes. On 524 held-out NIFTY calls, validation-selected test-time adaptation reduces DeepONet RMSE by 28.3%, but per-expiry mixture and SVI fits remain much more accurate. The evidence supports target-dependent inductive bias, not a universal winner.
Lennon Jason Shikhman, Michael Galarnyk, Aadi Dash et al.· 0 citations
Detecting a change in a multivariate series answers only the first of two questions; the operational question is which coordinates changed. Existing answers are incomplete. Block-level procedures certify predefined groups of coordinates under an additive union bound, high-dimensional variable-selection methods return interpretable rankings without error guarantees, and the post-detection inference literature controls error along the time axis rather than across coordinates. We propose ARM (Attribution by Rank Maxima), a wrapper that accepts a changepoint located by an arbitrary detector and returns the set of coordinates certified to have changed, each carrying a location or scale type label. ARM scores each coordinate by a max-over-splits rank statistic. Because this statistic dominates the corresponding statistic at the estimated split, the resulting certificate is invariant to the manner, and to the accuracy, of the changepoint estimate. Three finite-sample guarantees follow from within-coordinate ranks alone: per-coordinate validity under any detector; exact family-wise error control through a Westfall--Young joint permutation that preserves cross-coordinate dependence, with a fully distribution-free Holm fallback; and false discovery rate control under arbitrary coordinate dependence in high dimensions through Benjamini--Yekutieli and e-BH. In simulations, naive per-coordinate testing at the estimated changepoint inflates its family-wise error beyond $0.66$ as the dimension grows, whereas ARM maintains the nominal level while retaining validity under heavy tails, power in high dimensions, and accurate type labels. On five financial series surrounding the 2008 collapse, ARM attributes a scale change to every asset class and excludes injected control coordinates.
Chenchen Peng, Mixia Wu, Qijing Yan et al.· 0 citations
We present in this article a non-parametric value-at-risk (VaR+CVaR) algorithm that remains accurate for an arbitrarily large number of underlying positions. The algorithm solves the two inherent problems of VaR estimation. First, past history is not directly applicable to the future, but all predictions of the future are based on the past. Second, VaR estimation is equivalent to modeling a single corner of a high-dimensional space (the corner where all bets lose simultaneously). The algorithm only uses mathematical methods that strictly do not degrade in accuracy at high-dimensions. Historical data are then directly incorporated with all high-dimensional relationships present, without manipulation. We test the algorithm with an ensemble of 500 portfolios with random positions across 49 distinct liquid futures of different expiries (VIX, equity indexes, gov. bonds, rates, energy, metals, livestock, agriculture, and softs). All VaR estimations are performed strictly blind to the future. The median portfolio rate of loss exceeding the 99% confidence daily VaR estimate is between $1.0\pm0.1$% depending on algorithm input parameters. 68% of portfolios have a rate of loss exceeding 99% VaR between $1.0\pm0.3$%, and 95% of portfolios between $1.0\pm0.5$%.
Kernel Stein Discrepancy (KSD) compares a sample to a fixed target distribution known only through its score, and is widely used for goodness-of-fit testing, sample quality assessment, and approximate inference. We study the estimation of $\operatorname{KSD}(P_0,P)$ from $n$ independent observations and identify the sharp spectral constant governing the minimax risk: it is the Hilbert-Schmidt norm of the Stein covariance operator $C_\star$, giving the minimax scale $\sqrt{\|C_\star\|_{\mathrm{HS}}/n}$. This scale is attained by the positive-part square-root U-statistic, whereas the standard plug-in V-statistic remains at the trace scale $\sqrt{\operatorname{tr}(C_\star)/n}$ and is therefore suboptimal by the fourth root of the effective rank of $C_\star$; for a Gaussian target with a fixed-bandwidth Gaussian kernel this factor is exponential in the dimension.
We ask whether any one-parameter structural correction to the radial acceleration relation (RAR) can be uniquely recovered from SPARC rotation curves, and answer with an identifiability audit: each candidate is benchmarked against per-galaxy nuisance freedom, with predictive scoring against mass-only and data-quality baselines. In the full sample (N = 126) the answer is no: a hybrid compactness term improves the fit, but zero-point freedom absorbs the gain, and in cross-validation the model fails to out-predict a mass-only baseline and loses to a quality-flag baseline. One regime retains structural information: in gas-dominated, low-acceleration disks -- where MOND's strict locality and $\Lambda$CDM feedback models diverge most sharply -- the RAR residual correlates with compactness ($r=0.46$, $p=1.3\times10^{-4}$), remains significant under hierarchical partial pooling ($\beta=0.23$, $p=1.7\times10^{-5}$; N = 63), and survives canonical joint control for quality, sampling, mass, inclination error, and first-order pressure support ($r=0.30$, $p=0.02$). All significant results pass a Benjamini-Hochberg correction over the declared 27-test family. Three limits temper that survival: it is not significant under rank-based control over the widest proxy set; it resides in faint dwarfs independent surveys do not reach; and after mass control it is shared across the mass-size manifold. Pressure support brackets the interpretation -- isotropic drift correction absorbs a quarter of the amplitude, while a Jeans treatment overcorrects resolved cases -- leaving the physical origin undetermined. The audit's product is the extraction limit: claimed corrections must clear the 0.106 dex per-galaxy nuisance floor, a mass-only baseline, and data-quality stratification.
Lukas A. Sosna· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.