Skip to content

Category

machine learning

1,973 papers

#machine learning Preprint Aug 2026

Federated Compositional Muon Optimizer for Matrix-Wise Models

This work proposes an effective federated compositional Muon (FedCoMuon) optimizer to solve distributed matrix-wise compositional optimization problems and proposes a variance reduced variant of FedCoMuon (FedCoMuon-VR) based on a momentum-based variance reduced technique.

Wang Yan, Feihu Huang · 0 citations
#machine learning Preprint Open access Aug 2026

Geometric and Behavioral Stratification in Transformer Residual Streams

Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream. But what kind of direction does such a basis select? We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined privileged anchor. Measured with respect to this anchor, residual-stream variation is geometrically and behaviorally stratified by proximity to the prediction. The stratification holds in all eighteen models tested (dense and mixture-of-experts, 7B-120B, base and instruction-tuned). A narrow, scale-invariant prediction interface concentrates readout-relevant structure, while the vast prediction-distal complement expands with model scale. Because the prediction direction sits nearly orthogonal to the principal variance axes, variance-based analyses recover this organization only partly, and the shortfall grows with prompt heterogeneity. Anchoring reveals a steep geometric gradient: prediction-proximal regions are highly structured and cluster related prompts, while the complement is flatter and anti-discriminates among prompt groups. The interface is a narrow slice but functionally decisive. Disrupting the variance directions closest to the prediction causes immediate divergence and frequent task-frame shifts; disrupting the next level down delays divergence and preserves framing. The complement is weakly readout-aligned per direction yet causally and temporally load-bearing, and behavior is driven by direction rather than magnitude. These results establish the prediction direction as a privileged anchor distinct from previously described coordinate axes, and give a geometric account of how high-dimensional computation coexists with linear readout.

Nelson Guda · 0 citations
#machine learning Preprint Aug 2026

Online Learning of Scale Parameters in Score-Driven Filters

The central observation is that the negative product-of-scores feedback employed in accelerated score-driven recursions can be read as the stochastic gradient of this predictive loss, offering a new variational perspective.

Fabrizio Lillo, Giulia Livieri, Gianluca Palmari · 0 citations
#machine learning Preprint Aug 2026

A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

This study focuses on diagnosis-related features and compares five feature selection paradigms for opioid use disorder prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection.

Zihan Ding, Yinan Liu, Tengfei Ma et al. · 0 citations
#machine learning Preprint Aug 2026

A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics

Magnetic components in high-frequency, high-power-density converters are increasingly driven by non-sinusoidal flux-density waveforms with fast transitions, minor-loop operation, dc bias, and temperature variation. Under these conditions, steady-state core-loss formulas and single-valued material curves cannot fully capture transient magnetization responses. This work proposes the Physics-Informed Hybrid Neural Operator (PI-HNO), a compact material-specific neural model with B-H energy-consistency regularization for core-loss-oriented transient magnetization prediction. Given the measured B(t)-H(t) history, the input B(t) series over the prediction interval and operating-condition information, PI-HNO predicts the H(t) series and the corresponding reconstructed B-H trajectory. The model integrates a local recurrent branch for boundary-state representation and rate-dependent response evolution with a Preisach-inspired global branch that extracts waveform-level hysteresis context. Evaluation on the MagNetX transient database using material-specific models for 14 ferrite materials demonstrates that PI-HNO achieves a compact trade-off between sequence accuracy and B(t)-H(t) energy consistency, with the mean and 95th percentile B(t)-H(t) energy consistency errors of 1.92% and 7.60%, respectively, using only 4777 trainable parameters per model. Ablation studies further demonstrate that the local, global, and energy-aware regularized components provide distinct contributions to transient magnetization prediction.

Yachao Zhu, Qiujie Huang, Sinan Li et al. · 0 citations
#machine learning Preprint Jul 2026

ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling

This work investigates the expressiveness of general query/key rotations and finds that any normalized continuous positive-definite attention modulation function can be approximated by random rotations induced by its own Fourier transform, which is term Random Fourier Rotations.

Yiwen Chen, J. Ainslie, Krzysztof Choromanski et al. · 0 citations
#machine learning Preprint Jul 2026

GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks

GEqTrain is presented, a configuration-driven framework that separates dataset semantics, model composition, and training objectives, and GEqDiff, a generative extension based on equivariant flow matching that aims to make equivariant modeling more reproducible, extensible, and reusable.

Daniele Angioletti, Marco Nobile, V. Limongelli · 0 citations
#machine learning Preprint Jul 2026

SEE: Structure-aware Exploring&Exploiting for Long-horizon GUI Agent Trajectory Synthesis

See, a two-stage data synthesis framework consisting of an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling, yields reproducible and explainable data generation.

Zhuohang Fan, Beichen Zhang, Yuanfa Li et al. · 0 citations

Low-dimensional topology of deep neural networks

The results suggest that low-dimensional topology can be a useful tool to guide designs of AI architectures, and generalize the results from $d = 3$ to arbitrary $d>3$.

Jun Ren, Lek-Heng Lim · 0 citations

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

It is shown that the asymptotic decay behavior of f is not fixed by the architecture and emerges from the coupling between the state dynamics and parameter dynamics, settling into either a collapsed regime (fast, exponential forgetting) or an extended, anti-collapsed regime (slow, power-law forgetting).

L. Livi · 0 citations

How Transparent is DiffusionGemma?

A suite of interpretability case studies are conducted, uncovering initial evidence of novel diffusion-specific phenomena such as non-chronological reasoning, token and sequence smearing, and intermediate-context reasoning, and monitorability is found, which finds that DiffusionGemma is similarly monitorable to Gemma 4.

Joshua Engels, C. McDougall, Bilal Chughtai et al. · 1 citation

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.