Movie Soundtrack Emotion Matching Algorithm Based on CLIP Model
Abstract
To address the semantic distortion and temporal misalignment introduced by existing movie soundtrack matching methods that rely on textual intermediaries, this paper proposes an end-to-end audio-visual cross-modal temporal alignment algorithm based on the CLIP model. Considering that robust multimodal signal fusion and temporal synchronization are increasingly important for intelligent information processing in advanced electromagnetic sensing and communication environments, the proposed framework constructs a unified embedding space without textual bridging by extracting visual semantics through CLIP-ViT-B/32 and integrating AudioMAE with BiLSTM to capture deep temporal emotional characteristics of music. Musical representations are directly aligned with visual emotion prototypes to reduce the semantic gap, while a sliding-window temporal attention mechanism enables fine-grained dynamic soft alignment between scene transitions and musical emotional evolution. Furthermore, a multi-dimensional weighted scoring function incorporating semantic consistency, rhythm synchronization, and emotional intensity is designed, and global recommendation sequences are optimized under scene continuity constraints. Experimental results demonstrate that the proposed method achieves an emotion category accuracy of 76.3 ± 0.6%, an average cross-modal similarity of 0.732 ± 0.088, and an average temporal alignment deviation of only 1.24 ± 0.22 s. The framework significantly improves temporal synergy and emotional consistency between visual and audio modalities, providing an effective solution for cross-modal signal alignment and offering valuable methodological insights for multimodal information fusion and intelligent signal processing in electromagnetic wave perception and communication-oriented applications.