Quality-Aware Multimodal Graph Learning Via Attention-Guided Fusion and Reliability-Aware Optimization
Abstract
Multimodal graph learning has become an important direction for modeling graph-structured data with heterogeneous node content, such as text and images. However, existing methods often rely on globally shared or weakly adaptive fusion strategies, which makes it difficult to model node-wise differences in modality usefulness under graph context. In addition, they usually treat all nodes and edges as equally reliable during training, even though real-world multimodal graphs may contain weak cross-modal agreement, limited structural support, and noisy higher-order neighborhoods. To address these issues, we propose a quality-aware multimodal graph learning framework that integrates modality-specific projection, node-level attentionguided fusion, graph context encoding, and reliability-aware optimization. First, we design a node-level attention-guided fusion module with a stabilizing averaging branch, which adaptively adjusts modality contributions for each node while preserving shared cross-modal information. Second, we introduce a nodewise quality prior constructed from multimodal consensus, local semantic consistency, structural support, and second-order neighborhood coherence. This quality prior is further injected into downstream learning as a reliability-aware training signal for both node classification and link prediction. In this way, the proposed framework jointly models multimodal interaction and node reliability within a unified pipeline. Experiments on five MM-GRAPH benchmarks show that the proposed method achieves the best performance on all evaluated datasets under the adopted feature setting. For example, our method reaches 88.00% and 87.41% ACC on Ele-fashion and Goodreads-NC, respectively, and overall performs best among the compared methods across both node classification and link prediction tasks. Our code is publicly available at https://github.com/salotin/QA_MM_Graph.