Multimodal Fusion of Financial Text and Time Series Based on Contrastive Learning for Price Prediction
Abstract
Financial texts such as company-specific news and earnings reports contain useful signals about the future stock price movements of individual stocks, and their use in stock price prediction has gained growing attention. A common approach is to encode both financial text and historical price time series into embeddings and use them jointly as inputs to prediction models. However, when encoders are trained only on a downstream prediction objective, the text and time-series embeddings may not be well aligned. This misalignment may hinder the ability of prediction models to integrate the two modalities and can cause over-reliance on a single modality, limiting predictive performance. To address this limitation, aligning financial text embeddings and stock price time-series embeddings remains a promising yet underexplored direction. This study proposes a contrastive learning framework to align text embeddings of company-specific news with time-series embeddings of subsequent price sequences of individual stocks. After pretraining with this framework, the prediction model uses news text embeddings and past stock price time-series embeddings produced by the pretrained encoders as inputs. Because these embeddings are aligned, the prediction model can more effectively integrate information from both modalities in downstream stock prediction tasks such as stock movement direction classification and time-series forecasting. We evaluate the proposed method on binary stock movement direction classification, using a news–price paired dataset from the US stock market. Experimental results demonstrate that the proposed method achieves better classification performance than all comparison methods, including those without contrastive pretraining and those that use only a single modality as input. Furthermore, trading simulations show that the proposed method improves risk-adjusted trading performance, particularly in terms of a higher Sharpe ratio and a lower maximum drawdown.