This work evaluates ICL across 20 large language models on three antibody tasks: species-origin, antibody specificity, and isotype class classification, and introduces a sequence similarity-based strategy for ICL in antibody sequence classification, Sim-ICL.
Abstract
Large language models can learn new tasks through in-context learning (ICL), yet this ability remains underexplored for biological sequence classification. We evaluate ICL across 20 large language models on three antibody tasks: species-origin, antibody specificity, and isotype class classification. Few-shot prompting improves over zero-shot performance, but matching the performance of protein language model classifiers requires sequence-similar demonstrations. Building on this observation, we introduce a sequence similarity-based strategy for ICL in antibody sequence classification, Sim-ICL. Using 32-shot prompting, Sim-ICL achieves competitive performance on two of three tasks. Its simplicity makes few-shot ICL promising for antibody characterization, especially for researchers with limited coding expertise.
It is shown that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification, and MPEP conditioning is established as a lightweight strategy for low-data peptide classification.
Joshua Almonte, M. Vu, Andrew Ahn et al.· bioRxiv· 0 citations
This study presents a systematic empirical investigation of task-adaptive continual pre-training (TAPT), introduced by Gururangan et al., for Turkish language understanding, with a particular focus on the effect of the masked-language-modeling rate.
Murat Aydoğan, Savas Yildirim, Tuǧba Dalyan· IEEE Access· 0 citations
The results show that hybrid ensemble methods can improve token-level accuracy in low-resource POS tagging, while also revealing a trade-off between frequent-tag accuracy and rare-tag robustness.
Abdelouahed Moussaoui, Nor-Eddine Azalmad, Said Bahassine et al.· Information· 0 citations
Entity Resolution (ER) identifies and links records that refer to the same real-world entity. Rule-based approaches rely on explicit similarity functions, while deep learning and pre-trained language model (PLM)-based methods require large amounts of task-specific labeled data, both of which are often difficult to ob...
Hao-Yu Wang, Hai-Tong Tang, Jia-Jie Fu et al.· Proceedings of the VLDB Endo...· 0 citations
This work presents PromptSpLiCE, a post-hoc method that expresses each class-conditioned text embedding as a sparse combination of concepts from a fixed natural-language dictionary, using the same dictionary before and after prompt learning to compare changes in their concept profiles.
Ryousuke Kamiya, Hiroshi Kera, Kazuhiko Kawamoto· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.