GRASSP provides a competitive framework for integrating pretrained RNA representations with spatial structural context while reducing reliance on additional handcrafted structural annotations, and is demonstrated to outperform state-of-the-art baselines.
Abstract
MOTIVATION
RNA-small molecule binding site prediction is crucial for targeted drug discovery. Sequence-based methods are efficient but often fail to capture structural dependencies between nucleotides, whereas structure-aware graph models can better represent spatial interactions but typically rely on complex structural annotations and multi-stage preprocessing pipelines. We therefore developed GRASSP, a streamlined hybrid deep learning framework that integrates pretrained RNA language model (LM) representations with adaptive graph refinement.
RESULTS
GRASSP leverages nucleotide embeddings and predicted secondary-structure features from a pretrained RNA LM to construct spatial RNA graphs, followed by a lightweight two-step graph attention refinement module with adaptive gating to capture local and contextual nucleotide dependencies. Across four benchmark datasets (TE18, HARIBOSS, TL12, and JL10), GRASSP generally outperformed state-of-the-art baselines, with improvements of up to 24.1% in AUC and 44.5% in MCC. Ablation analyses showed that pretrained RNA representations provided the dominant predictive contribution, while spatial graph refinement offered complementary but dataset-dependent benefits. These results demonstrate that GRASSP provides a competitive framework for integrating pretrained RNA representations with spatial structural context while reducing reliance on additional handcrafted structural annotations.
AVAILABILITY
Code and datasets are publicly available at https://github.com/langiocn/GRASSP, with an archival snapshot available on Zenodo at https://doi.org/10.5281/zenodo.21888291.
SUPPLEMENTARY INFORMATION
Supplementary data are available at Bioinformatics online.
RNA-binding proteins (RBPs) orchestrate a complex combinatorial regulatory “code” that governs RNA splicing, stability, localization, and translation. Learning the relationship between RNA sequences and these processes is a central challenge in genomics. Foundation models, notably RNA language models, have emerged as the dominant approach, learning general-purpose representations from unlabeled sequence at scale. While RNA language models have demonstrated impressive performance across a broad range of downstream tasks, they generally learn from sequence reconstruction objectives alone, lacking direct connections to the regulatory principles that govern RNA function. Here we introduce Parnet, an RNA foundation model trained directly and exclusively on experimental CLIP-seq data. Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RBPs to predict base-resolution RBP binding profiles directly from RNA sequence. This CLIP-seq pretraining strategy departs fundamentally from the masked-language-modeling paradigm, anchoring learned RNA representations directly in measured protein–RNA interactions rather than sequence statistics. Parnet substantially outperforms its single-task predecessor RBPNet in binding profile and motif recovery, generalizes to unseen cell types and iCLIP data, and recapitulates position-dependent splicing regulation. Frozen Parnet embeddings, without task-specific fine-tuning, match or exceed the performance of both task-specific tools, as well as larger self-supervised RNA and genomic language models across diverse downstream tasks, including RNA biotype classification, lncRNA chromatin localization, translational efficiency, splice-site recognition, intron retention, and non-coding variant effect prediction. Importantly, Parnet remains mechanistically interpretable, tracing predictions back to the specific RBPs and motifs that drive them. These results establish the RBP interactome as a compact, functionally sufficient, and interpretable basis for foundation model pretraining in RNA biology.
Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural blindness because sequence representations under-specify key graph-topological cues. We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hint with confidence, and a GNN extracts an instance-specific explanatory subgraph (e.g., a subgraph SMILES and an accompanying explanatory paragraph). We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand. Across two datasets, enriching prompts with graph-derived context yields substantial accuracy gains, often exceeding 25% relative improvement and up to 74% on Tox21. We further validate the functional relevance of the extracted motifs via a necessity-based edge-drop intervention. Despite the observed gains, a persistent gap remains to specialized GNN models, highlighting both the value and limits of text-conditioned reasoning for molecular structure.
K. Bougiatiotis, Dimitrios Kelesis, Georgios Paliouras· 0 citations
The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.
Junkai Wang, G. Luo, Yun-Song Yang et al.· Bioinformatics· 0 citations
Experiments show that GraESM-FuseDTA achieves competitive overall performance and consistent advantages in ranking-oriented and variance-explanation metrics across warm start, drug cold start, target cold start, and strict pair cold start settings.
RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are costly and time-consuming, Deep Learning (DL) methods provide an efficient computational alternative for RPI Prediction (RPIP). In particular, Graph Neural Networks (GNNs) are promising, as they naturally model RPI networks. However, existing GNN-based methods often rely on homogeneous graphs or predefined meta-paths, which limit their ability to handle data sparsity and to generalize to cold-start scenarios involving unknown molecules. To address these limitations, we propose Edge Generation-guided Relation-aware Learning (EGRL), a novel framework with several key components: implicit meta-path learning to capture relational semantics without handcrafted paths; a multi-relation-aware attention mechanism for adaptive fusion of interaction patterns; a graph generator that predicts potential ("soft") edges to support cold-start nodes; and a multi-feature fusion predictor for final interaction scoring. EGRL is jointly trained with a primary task loss and an auxiliary generator loss. Comprehensive evaluations on four benchmark datasets demonstrate that EGRL achieves competitive overall performance. More importantly, it exhibits superior generalization in cold-start settings, achieving an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.867 and an Area Under the Precision-Recall curve (AUPR) of 0.861 on unknown molecules, corresponding to improvements of 8.6% in AUROC and 5.0% in AUPR over prior state-of-the-art methods. The code will be released soon.
Danyu Li, Ling Zhou, Rubing Huang et al.· 0 citations
Introduction Existing methods for gene regulatory network (GRN) inference rely primarily on gene expression data alone or on lower-resolution bulk sequencing data. Despite recent advances in integrating chromatin accessibility and RNA sequencing, inferring GRNs from paired single-cell multi-omics data remains challenging due to noise, sparsity, and complex nonlinear regulatory relationships. Methods We present MultiCausGRN, a graph attention network (GAT)-based framework for GRN inference from paired scRNA-seq and scATAC-seq data. The model incorporates directed prior-guided graph attention learning to capture biologically grounded regulatory directionality by integrating curated directed regulatory edges into graph representation learning. MultiCausGRN performs supervised transcription factor–target link prediction using integrated multi-omics features within a two-layer graph attention architecture. Results On the human PBMC multi-omics dataset, prior knowledge integration improved predictive stability and achieved a mean test AUPRC of 0.743 ± 0.049 and a mean AUROC of 0.682 ± 0.026 across five independent random seeds. Discussion These results demonstrate that directed prior-guided graph learning can improve the robustness and biological interpretability of GRN inference in data-limited settings. MultiCausGRN is publicly available at: https://github.com/nrr-90/MultiCausGRN.
N. Alkhateeb, Mamoun A. Awad· Frontiers in Bioinformatics· 1 citation
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.