Multimodal sentiment analysis aims to extract affective signals from multiple modalities and perform feature analysis, modality fusion and sentiment prediction. However, in practical application scenarios, input modality information is highly prone to be missing due to various objective factors, which further causes the model to produce biased or even completely erroneous sentiment judgment results. To address this critical issue, a model named Modal Experts and Missing Prompt Generation (MEMPG) is proposed. The model employs a brand-new dual-residual modality expert architecture to integrate the knowledge of hybrid modality experts. This architecture can not only retain the original information but also preserve the contextual information learned by the attention mechanism, thus exhibiting better robustness. The model adopts a twostage training strategy: In the first stage, each modality expert branch is independently pre-trained to extract low-level basic features of a single modality and strengthen the feature extraction capability of modality experts. In the second stage, all input modality features are fed into all modality expert branches for processing, and adaptive weights are assigned to modality features via an adaptive gating mechanism to obtain fused modality features with stronger representation ability. Subsequently, the enhanced features are input into the missing modality prompt generation module, which incorporates prompt learning to guide the reconstruction of missing modalities. Extensive comparative experiments are conducted on two standard multimodal sentiment datasets, namely MOSI and MOSEI, under various common modality missing scenarios. Experimental results demonstrate that, compared with current mainstream baseline models, the proposed MEMPG model achieves superior sentiment prediction performance.
Shuai Liu, Xuan-Yu Wu· 2026 IEEE International Conf...· 0 citations
Zero-Shot Anomaly Detection (ZSAD) aims to accurately identify anomalous samples from unseen categories without relying on target class training data. In industrial quality inspection scenarios, collecting training samples for target defect categories is often impractical due to production constraints and data scarcity, and ZSAD methods can effectively address the challenge of reliable anomaly detection under limited data conditions. Recently, vision-language models have shown strong generalization and inherent zero-shot capabilities, greatly facilitating their wide application in zero-shot industrial anomaly detection tasks with competitive and reliable detection performance. However, they have critical practical limitations: insufficient attention to fine-grained image details and poor adaptability to the specific requirements of industrial anomaly detection tasks. To address these limitations, we propose DPRF-CLIP, a CLIP-based ZSAD framework. It uses a pre-trained ResNet network to extract fine-grained local image features, which are fused into CLIP’s visual encoder via a specially designed bidirectional cross-attention module. A feature enhancement module is also integrated to further strengthen the model’s ability in capturing fine-grained local visual patterns. Comprehensive experiments on real-world industrial benchmarks (MVTec AD, VisA) show DPRF-CLIP achieves competitive performances, validating its effectiveness and strong generalization in industrial anomaly detection.
Shuai Liu, Tian-Jiao Ma· 2026 IEEE International Conf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.