Skip to content

Author

Xiaochun Cao

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Noise-Induced Cross-Modal Information Interaction and Dual-Prompt Learning for Medical Image Segmentation

Accurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model’s ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks—including cell, lung infection, and polyp segmentation—demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP

Chao Huang, Peng Chen, Jie Wen et al. · 0 citations
Aug 2026

GlyphTSR: Text-Rich Scene Image Super-Resolution Beyond Glyph Priors

Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficult to simultaneously preserve holistic scene fidelity and detailed character structure. To address the problem, this paper proposes a unified framework GlyphTSR that leverages glyph priors to enhance local glyph restoration while ensuring global semantic consistency. Specifically, the character aware extractor (CAE) with text-spotting guidance and the glyph region enhancer (GRE) with dual-tower encoder are designed to mine glyph features from low-quality inputs. CAE implicitly extracts latent glyph-related semantics to facilitate glyph restoration. GRE explicitly utilizes glyph and position cues to upgrade feature representation. The glyph-guided super-resolution with glyph guider and text-imbalance responsive loss (TIRL) is proposed to enhance local text restoration while performing global scene super-resolution. Furthermore, a new TS-ISR benchmark is established to jointly evaluate models’ ability to restore faithful scenes and legible text. Experiments on the proposed benchmark and public datasets demonstrate that the proposed GlyphTSR achieves competitive performance, establishing a baseline for the emerging TS-ISR. The benchmark is publicly available at https://github.com/qyx596/tsisr-benchmark

Na Jiang, Yuxuan Qiu, Jia-Wei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.