Skip to content

Category

machine learning

3,367 papers

#machine learning Preprint Aug 2026

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable but cross-implementation FP8 GEMM shows a different signature: both the prevalence and the magnitude of differences grow with reduction depth, while the INT8 fraction stays at parts per million and within one spacing over a 64x range of K.

Teng-Ruei Chen · 0 citations
#machine learning Review Aug 2026

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

CoMedBench is introduced, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets.

Akanta Das, Farhad Al-Amin Dipto, Mrinmoy Sarkar Anto et al. · 0 citations
#machine learning Preprint Aug 2026

Federated Compositional Muon Optimizer for Matrix-Wise Models

This work proposes an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems and proposes a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique.

Wang Yan, Feihu Huang · 0 citations
#machine learning Preprint Open access Aug 2026

Geometric and Behavioral Stratification in Transformer Residual Streams

Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.

Nelson Guda · 0 citations
#machine learning Preprint Aug 2026

Online Learning of Scale Parameters in Score-Driven Filters

The central observation is that the negative product-of-scores feedback employed in accelerated score-driven recursions can be read as the stochastic gradient of this predictive loss, offering a new variational perspective.

Fabrizio Lillo, Giulia Livieri, Gianluca Palmari · 0 citations
#machine learning Preprint Aug 2026

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

This study focuses on diagnosis-related features and compares five feature selection paradigms for opioid use disorder prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection.

Zihan Ding, Yinan Liu, Tengfei Ma et al. · 0 citations
#machine learning Preprint Aug 2026

A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics

Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these conditions, steady-state core-loss formulas and single-valued material curves cannot fully capture transient magnetization responses. This work proposes the Physics-Informed Hybrid Neural Operator (PI-HNO), a compact material-specific neural model with B-H energy-consistency regularization for core-loss-oriented transient magnetization prediction. Given the measured B(t)-H(t) history, the input B(t) series over the prediction interval and operating-condition information, PI-HNO predicts the H(t) series and the corresponding reconstructed B-H trajectory. The model integrates a local recurrent branch for boundary-state representation and rate-dependent response evolution with a Preisach-inspired global branch that extracts waveform-level hysteresis context. Evaluation on the MagNetX transient database using material-specific models for 14 ferrite materials demonstrates that PI-HNO achieves a compact trade-off between sequence accuracy and B(t)-H(t) energy consistency, with the mean and 95th percentile B(t)-H(t) energy consistency errors of 1.92% and 7.60%, respectively, using only 4777 trainable parameters per model. Ablation studies further demonstrate that the local, global, and energy-aware regularized components provide distinct contributions to transient magnetization prediction.

Yachao Zhu, Qiujie Huang, Sinan Li et al. · 0 citations
#machine learning Preprint Jul 2026

ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

This work investigates the expressiveness of general query/key rotations and finds that any normalized continuous positive-definite attention modulation function can be approximated by random rotations induced by its own Fourier transform, which is term Random Fourier Rotations.

Yiwen Chen, J. Ainslie, Krzysztof Choromanski et al. · 0 citations
#machine learning Preprint Jul 2026

GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks

GEqTrain is presented, a configuration-driven framework that separates dataset semantics, model composition, and training objectives, and GEqDiff, a generative extension based on equivariant flow matching that aims to make equivariant modeling more reproducible, extensible, and reusable.

Daniele Angioletti, Marco Nobile, V. Limongelli · 0 citations
#machine learning Preprint Jul 2026

SEE: Structure-aware Exploring&Exploiting for Long-horizon GUI Agent Trajectory Synthesis

See, a two-stage data synthesis framework consisting of an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling, yields reproducible and explainable data generation.

Zhuohang Fan, Beichen Zhang, Yuanfa Li et al. · 0 citations

Low-dimensional topology of deep neural networks

The results suggest that low-dimensional topology can be a useful tool to guide designs of AI architectures, and generalize the results from $d = 3$ to arbitrary $d>3$.

Jun Ren, Lek-Heng Lim · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.