Skip to content
Book Open access

Capturing Motif Topological Diversity via Geometry-Adaptive Riemannian Molecular Representation Learning

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 11444-11455 · 0 citations · 36 references

Abstract

Molecular properties are often governed by a small number of local substructures, or motifs, whose topologies can vary drastically across molecules. Existing molecular representation learning approaches typically embed all motifs into a single Euclidean or fixed-curvature space, which fails to capture the motif-level topological heterogeneity and leads to geometric mismatch, impairing property prediction. To address this challenge, we propose a geometry-adaptive Riemannian framework for molecular representation learning, which explicitly models motifs as the basic units and learns their embeddings across multiple constant-curvature spaces. Each motif is adaptively aligned with the geometric space that best fits its intrinsic topology, enabling simultaneous modeling of cyclic, hierarchical, and tree-like structures. Motif embeddings are then aggregated into molecule-level representations, emphasizing functional substructures while suppressing irrelevant background. Extensive experiments on benchmark molecular property prediction datasets demonstrate that our approach outperforms state-of-the-art baselines, shows strong generalization under distribution shifts, and provides interpretable motif-level insights, offering a general and scalable framework for scientific molecular modeling. Our code is available at https://github.com/qimuya/mo-mi-r.

Read PDF

Similar papers

#machine learning Preprint Aug 2026

Structural Hierarchy and Geometry in Molecular Representation Learning

Results show that explicitly teaching the relation between a molecule and its structural core can reliably shape the organization of molecular embedding space, while the extent of usefulness of this organization remains task dependent.

David Sulu, Lorenzo Di Fruscia, Jana M. Weber · 0 citations
Open access Aug 2026

OmniScore: Universal Scoring of Diverse Biomolecular Complexes via Equivariant Geometry-Aware Discrete Representation Learning

Scoring biomolecular complexes is central to structure assessment and drug discovery, yet the complexes themselves vary widely in pose, size, and molecular composition. A scoring function tuned for one interaction type rarely carries over to another, and most existing methods compound the problem by leaning heavily on task-specific labels. We introduce OmniScore, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads. OmniScore couples a graph view and a sequence view of each structure, encodes its three-dimensional geometry, and compresses representations into a compact latent space that a reconstruction module and prediction heads can reuse. We pretrain this backbone on diverse datasets including complexes, monomers, and small molecules with complementary objectives: coordinate recovery, correcting corrupted input tokens, predicting molecular identity, and grounding the representation in structure-level physical quantities. Across the evaluated benchmarks, OmniScore gave the best antibody-antigen and nanobody-antigen quality assessment on all reported metrics compared to state-of-the-art baselines. Its frozen residue embeddings matched the state-of-the-art protein-tokenization method with an average functional-site accuracy of 71.8% on a standard residue-level benchmark. On protein-ligand scoring and ranking benchmarks, it performed on par with methods built specifically for that single task. These results suggest that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend on interfacial and residue-level structure, within the evaluated settings.

Tien-Cuong Bui, Junsu Ko, Ju-Yong Lee · 0 citations
Preprint Jul 2026

Reconstructing local environments from concise atomistic representations

This work investigates the inverse problem of recovering atomic structures from local invariant descriptors, and shows that accurate reconstructions can be obtained from remarkably compact descriptors of different correlation orders, each comprising only a few tens of features.

Jigyasa Nigam, T. Phung, Ameya Daigavane et al. · 1 citation
Jul 2026

Persistent Manifold Learning of Protein Properties

PML is introduced, a novel computational framework that describes a binding interface as a family of multiscale manifolds, and results indicate that much of what determines binding strength is encoded in the shape of the interface itself, and that a single geometric description serves both classes without hand-tailored features.

Xingjian Xu, Zhe Su, Guo-Wei Wei et al. · 0 citations
#machine learning Preprint Sep 2026

Topology-induced Operators Reveal Complementary Graph Representations without Training

Graph representation learning has largely focused on designing increasingly sophisticated models to transform graph topology into vector representations, or embeddings. However, the extent to which embedding quality depends on model learning, rather than on the underlying topological transformations, remains unclear. Here, we show that informative embeddings can be derived without complicated model design and gradient-based training. Propagating random features through implicit hierarchical structures induced by random walks and anonymous walks yields embeddings that capture node proximity and structural role, respectively. These two training-free embeddings preserve complementary aspects of graph organization and perform competitively with classic and recent methods across various node-, edge-, and graph-level tasks. They often require substantially less computation, resulting in a favorable quality-efficiency trade-off. Combining the two types of embeddings further improves inference quality of some tasks compared with using either embedding type alone. Our results suggest that informative graph embeddings can arise from carefully chosen topological transformations before any learning operation is applied.

Meng Qin, Jin-Qiang Cui, Hongwei Zheng et al. · 0 citations
Open access Aug 2026

Hierarchical discrete representations for coarse-to-fine protein conformation generation

This work proposes a novel approach that learns hierarchical discrete representations of protein structures using vector quantization, and outperforms state-of-the-art models such as ESMDiff across challenging benchmark datasets, including BPTI MD trajectories and conformational-changing pairs.

Seokjun On, Yujin Jeong, Kanghyeon Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.