Skip to content

Soft Non-diagonality Penalty Enables Latent Space-Level Interpretability of Parameter-Efficient Peptide LM at No Performance Cost.

Jul 2026 · Journal of Chemical Information and Modeling · Vol 66 16, pp. 9847-9857 · 0 citations · 35 references
Medicine

TL;DR

This work introduces a soft weight matrix nondiagonality penalty and a one-hot encoded sequence clustering-based contrastive loss to develop compact models with interpretable embeddings while maintaining competitive performance, and demonstrates that domain-tailored architectural design can yield highly parameter-efficient models with fast inference and preserved generalization capabilities.

Abstract

Emergence of large scale protein language models (pLMs) has led to significant performance gains in predictive protein modeling. However, it comes at a high price of interpretability, and efforts to push representation learning toward explainable feature spaces remain scarce. The prevailing use of domain-agnostic and sparse encodings in such models fosters a perception that developing both parameter-efficient and generalizable models in a low-data regime is not feasible. In this work, we explore an alternative approach to develop compact models with interpretable embeddings while maintaining competitive performance. With the bidirectional long short-term memory autoencoder (BiLSTM-AE) model trained on positional property matrices, we introduce a soft weight matrix nondiagonality penalty and a one-hot encoded sequence clustering-based contrastive loss. As evidenced by Jacobian analysis, the penalty aligns embeddings with the initial feature space, whereas the contrastive loss organizes the latent space semantically. This combination leads to consistent improvements in performance on a suite of eight common peptide biological activity and physicochemical properties benchmarks. The use of amino acid physicochemical properties and density functional theory (DFT) derived cofactor interaction energies as input features provides a foundation for intrinsic interpretability, which we demonstrate on fundamental peptide properties. The resulting model is over 33,000 times more compact than the state-of-the-art pLM ProtT5. It demonstrates performance stability across diverse benchmarks without task-specific fine-tuning, showcasing that domain-tailored architectural design can yield highly parameter-efficient models with fast inference and preserved generalization capabilities.

View source

Similar papers

Open access Aug 2026

Self-Attention over Parallel Dense Embeddings for High-Dimensional Omic Data

Results suggest that PLAT provides an effective and interpretable framework for high-dimensional transcriptomic classification and functional enrichment analyses consistently highlighted biological processes and disease pathways associated with breast cancer, supporting the biological relevance of the learned latent re...

Kamal Elatifi, Nicolas Jäger Gallego, Á. Sánchez-Pla et al. · 0 citations
Preprint Aug 2026

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

It is found that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure.

Nikolai Bolik, Lennart Stöpler, Artur Andrzejak · 0 citations
#machine learning Preprint Sep 2026

PACE: Plug-and-Play Contextual Embedding for Feature Screening with Pretrained Tabular Foundation Models

PACE (Plug-and-Play Contextual Embedding) inserts a frozen tabular foundation model (TFM) column encoder before an existing feature-scoring rule, expanding each feature into a higher-dimensional contextual representation, positioning pretrained column geometry as a reusable upstream primitive for tabular learning.

Qi Qin, Er-Bo Li, Ting Wei et al. · 0 citations
#machine learning Preprint Sep 2026

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.

Sankar Behera, D. Singh, Anshika Agnihotri et al. · 0 citations
Open access Sep 2026

TeZO: Empowering the Low-Rankness on the Temporal Dimension in the Zeroth-Order Optimization for Fine-Tuning LLMs.

Zeroth-order optimization (ZO) has demonstrated remarkable promise in efficient fine-tuning tasks for Large Language Models (LLMs). In particular, recent advances incorporate the low-rankness of gradients, introducing low-rank ZO estimators to further reduce GPU memory consumption. However, most existing works focus so...

Y. Sun, Qi-Xin Zhang, Liang Ding et al. · 11 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.