Skip to content
Open access

Bridging linguistic reasoning and biophysical reality toward peptide engineering via instruction-tuned language modeling

Aug 2026 · Bioinformatics · Vol 42 · 0 citations · 34 references
Medicine

TL;DR

Pep-Instructions is established as a unified benchmark and the value of peptide-specific instruction tuning for peptide understanding, prediction, and design is demonstrated.

Abstract

Abstract Motivation Peptides serve as critical mediators in biological systems, regulating essential processes ranging from neurotransmission to immune response. However, their nonlinear sequence–function relationships and immense chemical diversity pose significant challenges for efficient experimental characterization and therapeutic development. While Protein Language Models have advanced biological sequence understanding, they predominantly capture global evolutionary features of full-length proteins, often overlooking the local physicochemical dependencies and short-range residue interactions that are essential for defining peptide bioactivity. Results To address this gap, we leverage the linguistic competence of general-purpose Large Language Models (LLMs) to treat amino acid sequences as “biological text,” bridging natural language supervision with biochemical sequence modeling without relying on explicit structural or evolutionary priors. We curate a peptide-specific instruction dataset, Pep-Instructions, spanning function description, sequence design, property prediction, and physicochemical optimization, and adapt a general-purpose LLM through parameter-efficient instruction tuning. Extensive benchmarking against general-purpose language models shows consistent improvements across the evaluated peptide-centric tasks. In particular, the instruction-tuned model produces more semantically faithful functional descriptions, generates peptide sequences with stronger sequence-level similarity, with representative ESMFold case studies suggesting backbone-level consistency, improves prediction of diverse peptide properties, and enables more reliable directional optimization of physicochemical properties under the adopted in silico evaluation protocols. Overall, these results establish Pep-Instructions as a unified benchmark and demonstrate the value of peptide-specific instruction tuning for peptide understanding, prediction, and design. Availability and implementation Source code and Pep-Instructions are available at https://github.com/kjY7836/pepinstruction. Fine-tuned model weights are available at https://huggingface.co/Codelife176/Pep-instruction.

Read PDF

Similar papers

Open access Sep 2026

Evaluating Large‐Language Models in Bioinformatics Applications

Large language models (LLMs) have significantly revolutionized natural language processing through their strong capabilities in text generation and reasoning. Yet, their applicability to bioinformatics applications remains largely unexplored. Here, we systematically evaluate state‐of‐the‐art LLMs across six represent...

Hengchuang Yin, Zi-Wen Cui, Dong-Xu Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...

Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al. · 0 citations
Open access Sep 2026

Protein Language Models: Learning From Evolution, Designing Beyond It

Protein language models (PLMs) have transformed our ability to learn from evolutionary sequence space, but protein engineering ultimately asks a different question: not what evolution selected, but what we should build next. Zero-shot likelihoods therefore provide useful, but not universal, measures of fitness and can...

Maurice Brenner, Julius Schlensok, A. Plaikner et al. · 0 citations
#small language model Open access Aug 2026

Multi-Peptide Prompting Enables In-Context Learning in Protein Language Models

It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.

Joshua Almonte, M. Vu, Andrew Ahn et al. · 0 citations
#small language model Open access Sep 2026

Genolator enables protein function interpretation using a multimodal large language model fusing genomic and structural interpretation with natural language interaction

Genolator is presented, a multimodal large language model that integrates embeddings from DNA sequences, amino acid sequences, and protein structures with natural language queries and represents a step towards bridging genomic code and human language through the integration of a multimodal LLM.

M. Danner, Tanhim Islam, M. Begemann et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.