Sentiment analysis of multimodal social media data is of great importance, not only for recognizing objective information but also for capturing subjective emotional states. While single-modal sentiment analysis has achieved notable progress, existing multimodal approaches still face two key challenges: (1) inadequate modeling of subjective emotional characteristics and (2) insufficient handling of cross-modal inconsistencies. To address these limitations, we propose an image-text multimodal sentiment analysis framework (CogSent) grounded in the psychological Dual-System Theory. First, inspired by human fast and slow thinking, we develop a Hierarchical Dual-Channel Cognition (HDCC) architecture to extract intuitive and rational affective features, respectively. Second, we introduce an Intuition-Guided Cognition Refinement (IGCR) module that uses System-I-inspired intuitive representations as affective priors to retrieve and refine System-II-inspired contextual representations via cross-attention. Third, we propose a Dynamic Cross-Modal Cognition Fusion (DC2F) network that predicts sample-adaptive thresholds from modality discrepancy, agreement, and attention statistics, dynamically regulating directional image–text interaction to mitigate interference from conflicting cross-modal sentiment signals. Extensive experiments on public benchmarks demonstrate that CogSent achieves improved performance compared to existing methods and yields competitive results on multiple evaluation settings.
Guo-Guo Ye, Qi-Qi Chen, Li-Qi Yan et al.· Big Data and Cognitive Compu...· 0 citations
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the limited context windows and shallow representational capacities of lightweight text encoders when processing lengthy, terminology-dense clinical reports. While integrating medical large language models (LLMs) offers unprecedented clinical reasoning capabilities, it introduces three major bottlenecks: (i) the anisotropic representational collapse of generative LLMs under standard contrastive objectives, (ii) the prohibitive memory overhead of joint end-to-end training with large batch sizes, and (iii) the medical hallucinations induced by vanilla contrastive losses that ignore fine-grained anatomical laterality and negation modifiers. To address these challenges, we propose \textbf{SCALPEL}, a \textbf{S}emantic \textbf{C}ross-modal \textbf{A}lignment framework via \textbf{L}LM-\textbf{P}owered \textbf{E}ncoder \textbf{L}earning. First, Clinical Report Contrastive fine-tuning converts a generative LLM into an isotropic encoder via domain-specific clinical text adaptation. Second, an asymmetric alignment strategy leverages offline feature caching to enable efficient training. Critically, we formulate an Anatomy-Negation Aware Objective that explicitly penalizes mismatched image-text pairs involving laterality confusion or false negations. Extensive experiments across MIMIC-CXR, CheXpert, and IU X-Ray benchmarks demonstrate that SCALPEL achieves state-of-the-art performance in cross-modal retrieval, zero-shot disease classification and medical visual question answering.
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.