A feasibility study based on a 2.5D U-Net architecture to detect GME in space-time connected data, resulting in improved detection of moving GMEs against the background with respect to classical spot detection algorithms and 2D U-Net, yet retaining real-time execution speed with respect to more complex deep-learning architectures is proposed.
Andrea Angino, Ken Trotti, D. U. Pizzagalli et al.· 0 citations
A performance recovery framework based on Self-Distillation Fine-Tuning (SDFT) that effectively restores model capabilities and offers new insights into the internal mechanisms of self-distillation is introduced.
This work formalises LH in a core probabilistic programming language (PPL) and gives sufficient syntactic conditions for its prevention, proving that a safe language fragment satisfying these conditions cannot produce likelihood-hacking programs.
Jacek Karwowski, Y. Kaddar, Zihuiwen Ye et al.· arXiv.org· 2 citations
This work proposes a flexible, effective sampling method for masked language models (MLMs), and reports results from an extensive in vitro head-to-head evaluation for the antibody engineering setting, revealing that the choice of sampling method can have a substantial impact, motivating future research into this under-explored area.
Calvin McCarter, Nicholas Bhattacharya, Sebastian W. Ober et al.· arXiv.org· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
Genomic language models (GLMs) have emerged as powerful tools for learning representations of DNA sequences, enabling advances in variant prediction, regulatory element identification, and cross-task transfer learning. However, as these models are increasingly trained or fine-tuned on sensitive genomic cohorts, they risk memorizing specific sequences from their training data, raising serious concerns around privacy, data leakage, and regulatory compliance. Despite growing awareness of memorization risks in general-purpose language models, little systematic evaluation exists for these risks in the genomic domain, where data exhibit unique properties such as a fixed nucleotide alphabet, strong biological structure, and individual identifiability. We present a comprehensive, multi-vector privacy evaluation framework designed to quantify memorization risks in GLMs. Our approach integrates three complementary risk assessment methodologies: perplexity-based detection, canary sequence extraction, and membership inference. These are combined into a unified evaluation pipeline that produces a worst-case memorization risk score. To enable controlled evaluation, we plant canary sequences at varying repetition rates into both synthetic and real genomic datasets, allowing precise quantification of how repetition and training dynamics influence memorization. We evaluate our framework across multiple GLM architectures, examining the relationship between sequence repetition, model capacity, and memorization risk. Our results establish that GLMs exhibit measurable memorization and that the degree of memorization varies across architectures and training regimes. These findings reveal that no single attack vector captures the full scope of memorization risk, underscoring the need for multi-vector privacy auditing as a standard practice for genomic AI systems.
Alexander Nemecek, Wenbiao Li, Xiaoqian Jiang et al.· 0 citations
This paper introduces a Multimodal Mixture-of-Experts (MMoE) module as a lightweight plug-in to empower Transformer-based time series models for multimodal forecasting, eliminating the need for explicit representation-level alignment.
Jiafeng Lin, Yuxuan Wang, Huakun Luo et al.· arXiv.org· 1 citation
Feature-Community-guided DICE (FCom-DICE), a perturbation strategy built on DICE that rewires a set of structurally influential edges and adjusts node features to reduce the distinctiveness exploited by GNN message passing is introduced.
Dalyapraz Manatova, P. Moriano, L. J. Camp· 0 citations
This work decomposes infections into trend, seasonal, and residual components and uses these signals to drive continuous-time latent dynamics while jointly forecasting and inferring time-varying transmission, recovery, and immunity-loss rates.
Yiqi Su, R. Lee, J. Cui et al.· arXiv.org· 1 citation
Two methodologies for modelling aggregated supply and demand curves in the EPEX SPOT Day-Ahead market are proposed, emphasizing generative models as a way to recover distributional variability and a low-dimensional parametric representation that yields deterministic point forecasts.
The proposed Cluster Aggregated GAN framework is established, a hybrid generative approach that routes each appliance to a specialized branch based on its behavioral characteristics that consistently outperforms baseline methods across metrics measuring realism, diversity, and training stability.
Zikun Guo, A. Adedigba, Rammohan Mallipeddi· arXiv.org· 1 citation
A weighted Hilbert-space framework is developed and sufficient conditions under which the row-stochastic design converges faster even with a smaller spectral gap are derived, by using a Rayleigh-quotient and Loewner-order eigenvalue comparison.
Bing Liu, Boao Kong, Limin Lu et al.· arXiv.org· 0 citations
This work empirically verify that the weak formulation, with a proper choice of test function and integration domain, effectively filters noisy data and explains why a weak form loss function is analogous to fitting a model to filtered data and provides a practical way to parameterize the weak form.
Xuyang Li, J. Harlim, R. Maulik· arXiv.org· 1 citation· ⚡1
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.