Skip to content

Author

Ruochi Zhang

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Similarity‐Enhanced Representation Learning of Non‐Canonical Amino Acids for Therapeutic Peptide Modeling

ABSTRACT Peptides combine the favorable pharmacokinetics of small molecules with the high specificity of biologics, making them promising therapeutics. Incorporating non‐canonical amino acids (ncAAs) further enhances drug‐like properties, yet modeling remains challenging due to chemically modified residues and combinatorial sequence diversity. Here, we introduce SinCAA, a similarity‐enhanced pretraining framework specifically designed to encode ncAAs. The framework is built on the principle that amino acids with similar 3D conformations induce minimal perturbations to peptide properties. It jointly optimizes two complementary self‐supervised tasks: contrastive learning guided by a conformational similarity metric to capture functional relationships among ncAAs, and masked node reconstruction to encode the unique chemical identity of each ncAA. Built on a graph transformer backbone, this dual “relationship–identity” supervision enables SinCAA to learn robust atomic representations that generalize from individual ncAA building blocks to full‐length peptides. SinCAA exhibits strong zero‐shot performance in peptide property prediction and consistently outperforms state‐of‐the‐art pretrained models across diverse benchmarks. This framework provides an efficient and interpretable approach for in silico prediction and ranking of ncAA‐containing peptides, accelerating candidate screening in therapeutic peptide discovery.

Chen-Cheng Xu, Le-Song Wei, Jian-Min Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation

Token aggregation converts token-level representations into fixed-dimensional sample representations, but most pooling methods operate only in the original token space. We introduce Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling. FLaG represents the nonredundant rFFT spectrum through concatenated real and imaginary components, summarizes spectral tokens with learnable latent queries, derives a sample-conditioned channel gate, and reconstructs modulated token representations for downstream aggregation. We evaluate the same architecture across ESM2-based antimicrobial peptide (AMP) activity prediction, ResNet18 image classification on CIFAR-10 and CIFAR-100, and three RoBERTa-based language tasks. FLaG achieves the best macro-averaged Spearman correlation coefficient, RMSE, and Recall@50 across four AMP backbone-species settings and the highest top-1 accuracy on CIFAR 10. It also achieves the best mean results on five of seven language metrics, although mean pooling remains strongest on STSBenchmark. AMP-side mechanistic analyses reveal low-frequency prediction sensitivity across most encoder layers, with increased relative high-frequency sensitivity in the final layer, and pronounced peptide-specific positional responses. The residual gate broadly amplifies spectral channels while preserving the low-frequency-dominated energy profile, whereas latent cross-attention exhibits sample- and species-specific spectral allocation. Overall, FLaG provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and downstream task. Supplementary materials, source code, and data are available at https://www.healthinformaticslab.org/supp/ and https://github.com/Kewei2023/AMPCliff/tree/FLaG.

Kewei Li, Rong Zhang, Xuelin Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.