Similarity‐Enhanced Representation Learning of Non‐Canonical Amino Acids for Therapeutic Peptide Modeling
Abstract
ABSTRACT Peptides combine the favorable pharmacokinetics of small molecules with the high specificity of biologics, making them promising therapeutics. Incorporating non‐canonical amino acids (ncAAs) further enhances drug‐like properties, yet modeling remains challenging due to chemically modified residues and combinatorial sequence diversity. Here, we introduce SinCAA, a similarity‐enhanced pretraining framework specifically designed to encode ncAAs. The framework is built on the principle that amino acids with similar 3D conformations induce minimal perturbations to peptide properties. It jointly optimizes two complementary self‐supervised tasks: contrastive learning guided by a conformational similarity metric to capture functional relationships among ncAAs, and masked node reconstruction to encode the unique chemical identity of each ncAA. Built on a graph transformer backbone, this dual “relationship–identity” supervision enables SinCAA to learn robust atomic representations that generalize from individual ncAA building blocks to full‐length peptides. SinCAA exhibits strong zero‐shot performance in peptide property prediction and consistently outperforms state‐of‐the‐art pretrained models across diverse benchmarks. This framework provides an efficient and interpretable approach for in silico prediction and ranking of ncAA‐containing peptides, accelerating candidate screening in therapeutic peptide discovery.