Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

2026

TSformer: An Effective Two-Stage Transformer Framework for Underwater Image Enhancement

Underwater images often suffer from various complex degradations, hindering reliable visual measurement. However, most existing enhancement algorithms primarily depend on multiscale spatial features that lack global consistency constraints, while other methods incorporating frequency-domain information tend to suffer from high model complexity and may introduce extra frequency-domain noise, which negatively impacts restoration accuracy. To address these challenges, we propose a feasible two-stage underwater image enhancement (UIE) method based on an improved Transformer architecture. In particular, we first design a space-frequency dual-guided preliminary enhancement module, which is responsible for decoupling and enhancing global frequency-domain features, to suppress background noise and mitigate attention misalignment in the Transformer backbone. In addition, we design an efficient frequency-guided attention block (FGAB) that explicitly modulates the value features to reduce frequency-domain noise while employing downsampling and channel-splitting strategies to calculate global-context attention across channels with sharp computational complexity decline. Moreover, we propose a channel fusion block (CFB) that dynamically evaluates the importance of global information across different channels and adaptively fuses encoder and decoder features to suppress noise-dominant channels. Finally, extensive experiments are carried out on the six public benchmark datasets, and the results demonstrate the effectiveness and strong generalizability of the proposed method.

Xikang Xia, Wei Peng, Hou-Jun Wang et al. · 0 citations
Jul 2026

ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning

Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that depth can be recovered as a globally smooth field through iterative denoising, which does not explicitly reflect the piecewise and scale-dependent organization of scene geometry. In practice, geometric structure emerges progressively across spatial scales, where coarse layout, surfaces, and boundaries are constructed in a hierarchical manner. Motivated by this observation, we introduce ARDepth, which formulates depth estimation as structured auto-regressive generation. Instead of recovering depth through global refinement, ARDepth progressively constructs depth representations as spatial resolution increases. To support this generative process, we introduce Scale-Progressive Conditioning (SPC) to inject multi-scale visual features at each generation stage, and Semantic-Aware Guidance (SAG) to provide scene-level semantic priors that enhance global structural consistency. Together, these designs enable the model to capture fine-grained local details while maintaining coherent global geometry. Empirical results demonstrate that our approach achieves strong performance and produces structurally consistent depth predictions across scales, validating auto-regressive generation as a promising alternative paradigm for geometric modeling.

Zijie Wang, Wei Zhang, Wei-Ming Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.