Skip to content

Category

machine learning

2,173 papers

#artificial intelligence Preprint Open access Aug 2026

The Authenticity Gap in Human Evaluation

Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores. However, little consideration has been given as to whether this approach faithfully captures human preferences. Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators. These assumptions are often violated in practice, in which case annotator ratings cease to reflect their preferences. The most egregious violations come from using Likert scales, which provably reverse the direction of the true preference in certain cases. We suggest improvements to the standard protocol to make it more theoretically sound, but even in its improved form, it cannot be used to evaluate open-ended tasks like story generation. For the latter, we propose a new human evaluation protocol called $\textit{system-level probabilistic assessment}$ (SPA). When human evaluation of stories is done with SPA, we can recover the ordering of GPT-3 models by size, with statistically significant results. However, when human evaluation is done with the standard protocol, less than half of the expected preferences can be recovered (e.g., there is no significant difference between $\texttt{curie}$ and $\texttt{davinci}$, despite using a highly powered test).

Kawin Ethayarajh, Dan Jurafsky · 0 citations

Deep Learning Based on Generative Adversarial and Convolutional Neural Networks for Financial Time Series Predictions

This paper proposes the implementation of a generative adversarial network (GAN), which is composed by a bi-directional Long short-term memory (LSTM) and convolutional neural network(CNN) referred as Bi-L STM-CNN to generate synthetic data that agree with existing real financial data so the features of stocks with positive or negative trends can be retained to predict future trends of a stock.

Wilfredo Tovar · 8 citations
#machine learning Preprint Aug 2026

A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators

High-fidelity finite-element simulations can provide accurate numerical predictions for side-branch resonators, but large simulation datasets are expensive to generate and purely data-driven surrogates may become unreliable when simulation-labelled data are scarce. This study develops an analytical-prior learning framework that reuses a low-cost analytical model to improve data efficiency under limited high-fidelity simulation budgets. Two complementary routes are considered. When the analytical model remains available at inference, it is retained as an explicit baseline and the simulation data are used to learn only the analytical-to-simulation discrepancy. When a self-contained predictor is required, the analytical mapping is first distilled from abundant low-cost evaluations into a learned prior and then calibrated with the limited simulation data. The framework is evaluated on rectangular side-branch Helmholtz resonators using 86 simulation-labelled geometries and 8,998 non-overlapping analytical-only geometries. The analytical model achieved a mean absolute error (MAE) of 1.333 Hz. Direct support vector regression (SVR) achieved 3.375 Hz, while residual SVR reduced the MAE to 0.426 Hz. A direct multilayer perceptron (MLP) achieved 1.109 Hz, whereas analytical-prior pretraining reduced the error to 0.556 Hz with frozen-prior residual adaptation and 0.371 Hz with full-model fine-tuning. Across training budgets of 20 to 70 simulation-labelled cases, both analytical correction and analytical-prior pretraining consistently improved data efficiency relative to direct learning. These results show that analytical prior information can substantially improve high-fidelity prediction when simulation data are scarce, with explicit correction and prior distillation serving complementary deployment needs.

Jiaming Li · 0 citations
#artificial intelligence Preprint Aug 2026

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

OceanDepths is introduced, the first open, global, regridded AI-ready dataset that pairs satellite-derived sea surface temperature, sea surface salinity, and sea surface height products with co-located EN4 subsurface temperature and salinity profiles, complemented by matched GLORYS12 ocean reanalysis data to support comparisons or multi-stage learning.

Simon Donike, Ruben Cartuyvels, A. I. Ferola et al. · 0 citations
#artificial intelligence Preprint Aug 2026

NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption

Noise Corruption is introduced, a Noise Corruption-based explanation framework, which perturbs each message through matched-norm random-direction corruption while preserving the expected squared message norm, and NICE, a Noise Corruption-based explanation framework, which learns a Stochastic Restoration Boundary under NC-induced uncertainty, balancing target-prediction restoration against compactness.

Ziluowen Luo, Jun Yin, Ruochen Liu et al. · 0 citations
#machine learning Preprint Aug 2026

CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework

This work introduces CrevasseSeg, a framework for binary segmentation over the terminus of Borebreen, Svalbard, and releases CrevasseSeg to support label-efficient segmentation research in remote sensing.

Steve Wallace, William D. Harcourt, Richard Hann et al. · 0 citations
#machine learning Preprint Aug 2026

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable but cross-implementation FP8 GEMM shows a different signature: both the prevalence and the magnitude of differences grow with reduction depth, while the INT8 fraction stays at parts per million and within one spacing over a 64x range of K.

Teng-Ruei Chen · 0 citations
#machine learning Review Aug 2026

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

CoMedBench is introduced, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets.

Akanta Das, Farhad Al-Amin Dipto, Mrinmoy Sarkar Anto et al. · 0 citations
#machine learning Preprint Aug 2026

Federated Compositional Muon Optimizer for Matrix-Wise Models

This work proposes an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems and proposes a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique.

Wang Yan, Feihu Huang · 0 citations
#machine learning Preprint Open access Aug 2026

Geometric and Behavioral Stratification in Transformer Residual Streams

Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.

Nelson Guda · 0 citations
#machine learning Preprint Aug 2026

Online Learning of Scale Parameters in Score-Driven Filters

The central observation is that the negative product-of-scores feedback employed in accelerated score-driven recursions can be read as the stochastic gradient of this predictive loss, offering a new variational perspective.

Fabrizio Lillo, Giulia Livieri, Gianluca Palmari · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.