TSformer: An Effective Two-Stage Transformer Framework for Underwater Image Enhancement
Abstract
Underwater images often suffer from various complex degradations, hindering reliable visual measurement. However, most existing enhancement algorithms primarily depend on multiscale spatial features that lack global consistency constraints, while other methods incorporating frequency-domain information tend to suffer from high model complexity and may introduce extra frequency-domain noise, which negatively impacts restoration accuracy. To address these challenges, we propose a feasible two-stage underwater image enhancement (UIE) method based on an improved Transformer architecture. In particular, we first design a space-frequency dual-guided preliminary enhancement module, which is responsible for decoupling and enhancing global frequency-domain features, to suppress background noise and mitigate attention misalignment in the Transformer backbone. In addition, we design an efficient frequency-guided attention block (FGAB) that explicitly modulates the value features to reduce frequency-domain noise while employing downsampling and channel-splitting strategies to calculate global-context attention across channels with sharp computational complexity decline. Moreover, we propose a channel fusion block (CFB) that dynamically evaluates the importance of global information across different channels and adaptively fuses encoder and decoder features to suppress noise-dominant channels. Finally, extensive experiments are carried out on the six public benchmark datasets, and the results demonstrate the effectiveness and strong generalizability of the proposed method.