Sentiment analysis of multimodal social media data is of great importance, not only for recognizing objective information but also for capturing subjective emotional states. While single-modal sentiment analysis has achieved notable progress, existing multimodal approaches still face two key challenges: (1) inadequate modeling of subjective emotional characteristics and (2) insufficient handling of cross-modal inconsistencies. To address these limitations, we propose an image-text multimodal sentiment analysis framework (CogSent) grounded in the psychological Dual-System Theory. First, inspired by human fast and slow thinking, we develop a Hierarchical Dual-Channel Cognition (HDCC) architecture to extract intuitive and rational affective features, respectively. Second, we introduce an Intuition-Guided Cognition Refinement (IGCR) module that uses System-I-inspired intuitive representations as affective priors to retrieve and refine System-II-inspired contextual representations via cross-attention. Third, we propose a Dynamic Cross-Modal Cognition Fusion (DC2F) network that predicts sample-adaptive thresholds from modality discrepancy, agreement, and attention statistics, dynamically regulating directional image–text interaction to mitigate interference from conflicting cross-modal sentiment signals. Extensive experiments on public benchmarks demonstrate that CogSent achieves improved performance compared to existing methods and yields competitive results on multiple evaluation settings.
Guo-Guo Ye, Qi-Qi Chen, Li-Qi Yan et al.· Big Data and Cognitive Compu...· 0 citations
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.