Skip to content
Book Open access

Caduceus: MoE Foundation Models for Unifying Biological and Natural Language

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 12715-12726 · 1 citation · 74 references

TL;DR

This paper introduces Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language, and incorporates a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering.

Abstract

Multi-modality pre-training on protein sequences with textual descriptions has enabled general-purpose protein language models. However, as the property descriptions span heterogeneous domains, we observe a severe data interference phenomenon : distinct protein residues often target domain-specific annotations, revealing partially inconsistent functional mechanisms across sources, which substantially leads to degraded performance. This paper addresses this overlooked issue with a novel Mixture of Property-Guided LoRA Experts (MoPGLE) architecture, efficiently fusing knowledge across diverse domains. Concretely, we introduce Caduceus, a family of MoE-enhanced foundation models built with a hierarchical pre-training paradigm to jointly integrate biological and natural language. Employing a property-guided gating router that assigns domain-specific protein tokens to different experts, the dual-granularity alignment approach reconciles signals across diverse functional mechanisms. To extend generalization beyond particular tasks, we further incorporate a multi-task instruction tuning phase, enabling robust protein parsing and natural language question answering. Extensive experiments on 17 benchmarks demonstrate that øurapproach mitigates the intrinsic data interference and consistently delivers optimal performance. The instruction-tuned Caduceus-Instruct provides precise protein elucidation, significantly surpassing Galactica-30B, Evolla-10B, and BioMedGPT-7B. The code of this paper is publicly available at https://github.com/zju-ai4s/Caduceus.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

LEMON-ZEST: Evolution-Informed Tokenization for Efficient Protein Language Modeling

Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and structure-based tasks, yet the potential of tokenization remains underexploited. Unlike human language, proteins preserve structure despite extensive sequence variation a p...

Biswajit Banerjee, Claudia A. Carreno, Anton S. Petrov · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to...

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within these models. We address this gap by contributing two...

Darin Keng, Zhewei Sun · 0 citations
Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.

Clara Meister · 1 citation
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...

Clara Meister · 0 citations
#natural language process... Preprint Sep 2026

LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models

Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a classifier over pretrained-model activations to rank candidate...

Pavel Tikhonov, Elena Tutubalina, I. Oseledets et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.