Jul 2026· Journal of Chemical Information and Modeling· Vol 66 16, pp.
9847-9857
· 0 citations· 35 references
Medicine
TL;DR
This work introduces a soft weight matrix nondiagonality penalty and a one-hot encoded sequence clustering-based contrastive loss to develop compact models with interpretable embeddings while maintaining competitive performance, and demonstrates that domain-tailored architectural design can yield highly parameter-efficient models with fast inference and preserved generalization capabilities.
Abstract
Emergence of large scale protein language models (pLMs) has led to significant performance gains in predictive protein modeling. However, it comes at a high price of interpretability, and efforts to push representation learning toward explainable feature spaces remain scarce. The prevailing use of domain-agnostic and sparse encodings in such models fosters a perception that developing both parameter-efficient and generalizable models in a low-data regime is not feasible. In this work, we explore an alternative approach to develop compact models with interpretable embeddings while maintaining competitive performance. With the bidirectional long short-term memory autoencoder (BiLSTM-AE) model trained on positional property matrices, we introduce a soft weight matrix nondiagonality penalty and a one-hot encoded sequence clustering-based contrastive loss. As evidenced by Jacobian analysis, the penalty aligns embeddings with the initial feature space, whereas the contrastive loss organizes the latent space semantically. This combination leads to consistent improvements in performance on a suite of eight common peptide biological activity and physicochemical properties benchmarks. The use of amino acid physicochemical properties and density functional theory (DFT) derived cofactor interaction energies as input features provides a foundation for intrinsic interpretability, which we demonstrate on fundamental peptide properties. The resulting model is over 33,000 times more compact than the state-of-the-art pLM ProtT5. It demonstrates performance stability across diverse benchmarks without task-specific fine-tuning, showcasing that domain-tailored architectural design can yield highly parameter-efficient models with fast inference and preserved generalization capabilities.
The results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.
Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder et al.· 0 citations
Results suggest that PLAT provides an effective and interpretable framework for high-dimensional transcriptomic classification and functional enrichment analyses consistently highlighted biological processes and disease pathways associated with breast cancer, supporting the biological relevance of the learned latent re...
Kamal Elatifi, Nicolas Jäger Gallego, Á. Sánchez-Pla et al.· Applied Sciences· 0 citations
It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak· 0 citations
PACE (Plug-and-Play Contextual Embedding) inserts a frozen tabular foundation model (TFM) column encoder before an existing feature-scoring rule, expanding each feature into a higher-dimensional contextual representation, positioning pretrained column geometry as a reusable upstream primitive for tabular learning.
Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
Sankar Behera, D. Singh, Anshika Agnihotri et al.· 0 citations
Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs). In particular, recent advances incorporate the low-rankness of gradients, introducing low-rank ZO estimators to further reduce GPU memory consumption. However, most existing works focus so...
Y. Sun, Qi-Xin Zhang, Liang Ding et al.· IEEE Transactions on Pattern...· 11 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.