The results indicate that mixtures of M\"obius distributions provide interpretable, parsimonious models for the studied steady state density distributions, that match or outperform the alternatives.
Wanxin Gao, I. Nikolaidis, Janelle J. Harms· 2 citations
A Mixture of Multicenter Experts (MoME) framework to address AI bias in the medical domain without requiring data sharing across institutions is proposed and validated using a multimodal target volume delineation model for prostate cancer radiotherapy.
Yujin Oh, Sangjoon Park, Xiang Li et al.· 0 citations
This work revisits this design choice and proposes a sub-center modeling framework for speaker embeddings, which improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance in zero-shot voice conversion.
Ismail Rasim Ulgen, J. Hansen, Carlos Busso et al.· 0 citations
These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO, and study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM.
Yunpeng Ba, Zhi Zheng, Yue Xie et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work builds a test bench on 892 yield response curves from two long running UK experiments, and sweeps the nitrogen to grain price ratio to cover all price scenarios, finding that machine learning fails as a predictor and pays as a profit scored correction to standard advice.
A neural regression model (LitEm) that enables transductive knowledge graph embedding models to predict numerical attributes within knowledge graphs and a co-training framework that jointly trains state-of-the-art transductive knowledge graph embedding models with LitEm, which improves link prediction performance mainly for bilinear models and simultaneously enables them to predict numerical attributes.
Rupesh Sapkota, Louis Mozart Kamdem Teyou, Moshood Yekini et al.· 0 citations
ContourKV, a training-free allocator built from the dropped-mass statistic, wins $93$ of $160$ paired comparisons against that state of the art and loses $22$ at the byte count of the budget-enforcing baselines, and it ties the strongest of them.
This work introduces the cross-predictive JEPA (JEPA-x), which grounds latent dynamics in privileged physical trajectories, and shows that direct physical-state regression improves decodability without improving forecastability or control, indicating that the benefit comes from shaping latent dynamics rather than merely encoding physical variables.
Kehan Wen, Ziming Li, Siyuan Luo et al.· 0 citations
Frozen RNA-type evaluations show that RIBOSPAN learns state-of-the-art RNA representations, with a particularly clear advantage on long RNAs, and emerges as the strongest encoder-only RNA foundation model, achieving state-of-the-art performance in both full-transcript biological property prediction and zero-shot mutation-fitness modeling.
Ziyuan Wang, Bohao Tang, Fei Zhang et al.· 0 citations
Results show that architecture is a major source of task-relevant structure in TPC embeddings and should be treated explicitly when assessing representation learning and developing reusable detector models.
T. Wheeler, M. Kuchera, R. Ramanujan et al.· 0 citations
This work revisits client drift from a novel frequency-domain perspective and uncovers a critical Spectral Bias of Drift: inter-client gradient divergence is predominantly concentrated in low-frequency components which encode client-specific distributional shifts, while high-frequency components representing fine-grained features remain relatively consistent.
Liyang Yuan, Yibo Yang, Dandan Guo et al.· 0 citations
OBJECTIVE
Interpretable brain-computer interface classifiers that generalize across subjects without calibration remain an open challenge. We evaluated whether prototype-based cross-attention can provide competitive, inherently interpretable event-related potential (ERP) classification across diverse paradigms under deployment-compatible conditions.
APPROACH
We propose ERP-XTTN (ERP Cross-Attention), a cross-attention architecture that routes input electroencephalographic peaks to fixed difference-wave prototypes via query-key-only cross-attention with no value projection. Classification is based directly on prototype similarity and a separate measure of component amplitude, so that the prototype content contributes to every decision by construction. Prototypes are derived automatically from prominent extrema in the training-fold grand-average difference wave. We evaluated across three public sources (BNCI Horizon 2020, HRI Cursor, and ERP CORE) encompassing eight ERP components (ERN, LRP, ErrP, N170, P300, N2pc, MMN, N400). Evaluations used leave-one-subject-out (LOSO) cross-validation with causal filtering at a three-channel montage, compared against EEGNet, EEG-Deformer, ERP Prototypical Matching Net (EPMN), and xDAWN with Riemannian geometry (xDAWN+RG).
MAIN RESULTS
At three channels, the mean performance gap between the best baseline and ERP-XTTN was 0.025 area under the receiver operating characteristic curve (AUROC). Prototype interventions confirmed that decisions depend on the physiological content of the prototypes rather than on the routing attention pattern alone. False positives morphologically resembled true positives more than true negatives did across all datasets, indicating classification errors are neurophysiologically explicable.
SIGNIFICANCE
ERP-XTTN generalizes across diverse ERP morphologies under causal, calibration-free conditions, while retaining competitive performance and decisions that depend directly on physiological prototype content at a three-channel montage. Unlike post-hoc explanation methods for black-box models, the basis of each decision is directly observable in the trained model itself. To our knowledge, this is the first epoch-level LOSO benchmark on ERP CORE.
Charlotte Genevier Wyman, L. Hirshfield· Journal of Neural Engineerin...· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.