PromptArbiter: Degradation-Aware Infrared–Visible Image Fusion via Text-Conditioned Modality Arbitration and Multi-Scale Cross-Modal Interaction
Abstract
Infrared–visible image fusion aims to combine thermal saliency from infrared imagery with the structural detail preserved in the visible spectrum. In existing fusion methods, the two modalities are fused through fixed channel mixing, the same text embedding is reused across decoder scales, and explicit cross-modal interaction is confined to the deepest stage. These design choices weaken adaptability when the degradation severity is highly unbalanced across modalities. To address this issue, we propose PromptArbiter, a degradation-aware strategy. The proposed model introduces a text-guided fusion gate to rebalance modality contributions according to prompt semantics, level-specific text projection modules to inject scale-adaptive language cues, and lightweight cross-modal attention blocks at intermediate resolutions. A CLIP-based contrastive loss further provides semantic-level supervision. Experiments on public benchmark data show that PromptArbiter achieves improved fusion performance.