Skip to content

Category

machine learning

3,367 papers

#machine learning Preprint Aug 2026

More Data Cannot Break a Symmetry: Identifiability by Design

Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist. The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free. We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention. In colour, where candidate geometries have closed form, we show that the failure is structural: sixty-four times the restart budget leaves a symmetric design unmoved while an asymmetric set at the same N recovers every time. Discriminating representational models and recovering a correspondence are essentially uncorrelated objectives (r = -0.02 over 3,000 subsets). Choosing nine colours by this diagnostic alone, without consulting any learned representation, moves all 93 model representations away from the degenerate point and cuts catastrophic alignment failures from 75% to 2% with the models, the layers, N and the solver all held fixed. The same risk arises wherever a regular design meets its candidate geometry's isometry group, including evenly spaced orientations, tones, or motion directions, and the check costs one function call before data collection.

Jing-Lin Xu, Christopher Kanan · 0 citations
#machine learning Preprint Aug 2026

Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics

An evolving benchmark suite of challenging, natively-spherical PDE datasets including a modified Galewsky jet, anomalous chained turbulence, Cahn-Hilliard decomposition, spherical Riemann shocks, Held-Suarez dry atmospheric transport and global ocean dynamics are released.

T. Muser, Giovanni Abati, Ivan Dokmanic · 0 citations
#machine learning Preprint Aug 2026

When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging

This work reveals that the recent Muon optimizer as a mechanism that regulates this factor by construction tightens the interference bound for both CL and MM, positioning Muon as a principled optimizer-centric approach complementary to existing solutions.

Shan Liu, Yuehan Yin, Yinghuan Shi et al. · 0 citations
#machine learning Preprint Aug 2026

A Deeper Analysis of Block-Sparse Featurizers

This work proposes several architectural changes to the BSF, including a Tournament Top-K selection rule that significantly reduces feature splitting, and extends the block paradigm to the crosscoder.

Alexandru-Iulius Jerpelea, Amith Ananthram · 0 citations
#artificial intelligence Preprint Aug 2026

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

This work introduces TraceML, which pairs human and agent work on the same competitions under one version-level schema, and releases the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.

J. Yan, Weiwei Sun, Si-Jie Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as GRPO.

Wenze Lin, Jiale Zhao, Xitai Jiang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

A hybrid GAN-guided diffusion framework that uses a pretrained Wasserstein GAN with gradient penalty (WGAN-GP) as a feature prior for conditional diffusion-based image restoration that consistently improves the quality of both degraded and low-resolution images.

Saif Ahmed, Ashadullah Galib, S. R. R. Antu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

JuryProbe is introduced, an empirical consensus-risk diagnostic for reference-free factuality judge panels, paired with a calibration-based routing policy, which estimates consensus risk from a labeled calibration probe using false-negative-only (FN-only) judge correlation and false-consensus lift.

Tianxing Zhou, Ruixi Lin · 0 citations
#artificial intelligence Preprint Aug 2026

ED-CSP: Crystal Structure Prediction from Electron Diffraction

This work introduces ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED spot sets and establishes a benchmark for generative crystal structure prediction from sparse ED observations and provides a foundation for future transfer to experimental data.

Germain Poloudenny, Arnaud Demortière, Yael Fregier · 0 citations
#artificial intelligence Preprint Open access Aug 2026

Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction

Predicting how held-out target genes respond to CRISPRi perturbation, and whether such predictions transfer across biological screens, is hard to evaluate: a representation can be informative within one screen yet fail across screens, while endpoint definitions and design factors such as sampling depth differ between datasets. We evaluate a frozen Geneformer representation under a locked, pre-registered protocol, with heads and model selection frozen before test evaluation, external outcome labels withheld until final unblinding, and analysis-governing decisions fixed before the evaluations they govern. In-distribution on the Virtual Cell Challenge (VCC), the frozen representation carries measurable predictive information beyond a dimension-matched random-feature control (Delta R^2 = +0.1645, 95% CI [+0.1375, +0.1920]), satisfying the pre-registered informativeness gate required before interpreting transfer. It then fails zero-shot transfer on both external screens (Spearman rho = -0.139 and -0.267), lying below that control on each. Adding a predefined magnitude block improves the representation externally (Delta rho = +0.032 and +0.143) but does not rescue transfer: both remain negative. A pre-registered, count-adjusted max-response secondary is positively associated with the outcome on both screens; we report it as correlational and secondary, not as a recovered magnitude signal. Finally, the VCC endpoint is strongly sample-size associated: a count-only linear model reaches R^2 = +0.4325, versus +0.2589 for the four magnitude scalars; adding those scalars to cell count improves R^2 by only +0.0017, so much of the aggregate-magnitude signal overlaps with cell count. A locked evaluation thus surfaces a transfer failure and a sampling-depth entanglement that a less controlled evaluation could obscure.

Mehrdad Shoeibi, Niloofar Yousefi · 0 citations

Where Steering Signals Come From: Activation Source Selection in Activation Steering

Tail subtraction is introduced, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals, and suggests that steering depends on representations of what the model is about to do, not merely on what has already appeared.

Jiaran Ye, Lingxu Ran, Zijun Yao et al. · 2 citations
#artificial intelligence Preprint Jul 2026

On the Depth Scalability of Logic Gate Networks

Results indicate that scalable LGN depth requires both stable optimization and credit-preserving information access, and introduce Input-Anchored Logic Gate Networks (IALGN), in which each gate combines a private hidden spine with a direct input anchor.

Taegun An, Dohun Kim, Haebeom Lee et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.