PolyCLIP: A CLIP-Based Multimodal Framework for Polymer Property Prediction
Polymers are fundamental to modern materials science because their backbone chemistry, monomer composition, and chain architecture can be systematically tuned to achieve a virtually unlimited range of mechanical, thermal, electronic, and optical properties. To accelerate the discovery and design of polymeric materials, machine learning (ML) has emerged as a powerful tool for predicting structure–property relationships with high accuracy and efficiency. Inspired by the remarkable success of multimodal contrastive models, such as CLIP, which demonstrate that aligning text and visual representations in a shared embedding space enables strong generalization, we ask whether such multimodal alignment can benefit polymer property prediction. Despite the growing adoption of ML in polymer informatics, existing models suffer from unimodal feature dependence, computational overhead from large-scale pretraining, and limited generalization across diverse properties, with cross-modal approaches remaining largely unexplored. To address these limitations, we propose PolyCLIP, a CLIP-based multi-modal model that integrates textual (PSMILES) and visual (2D molecular structure image) representations for polymer property prediction. Molecular images for polymers associated with 15 target properties are generated from PSMILES representations using RDKit, and feature-level fusion of text and image embeddings allows the model to learn cross-modal chemical relationships. Furthermore, combined token-level SHAP attribution and visual self-attention analysis reveal strong cross-modal consistency, demonstrating that the model focuses on chemically meaningful substructures in a property-specific manner. PolyCLIP outperforms existing unimodal and multimodal models, achieving up to 13.32%, 3.48%, and 7.50% relative R2 improvements on thermophysical, electronic, and optical (refractive index) properties, respectively, while maintaining competitive performance in gas permeability predictions. These results demonstrate that CLIP embeddings effectively capture chemical information in polymers and position PolyCLIP as a simple, scalable, and interpretable basis for multimodal polymer informatics and data-efficient materials discovery.