Skip to content

Author

Meng-Huan Zhang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Sep 2026

Mitigating Textual Noise in Multimodal FGVC via Hierarchical Semantic Purification and Multi-Stage Alignment.

Fine-grained visual classification (FGVC) plays a crucial role in the realm of computer vision. Recently, multimodal FGVC methods, leveraging textual descriptions as semantic guidance, have gained considerable attention. However, current approaches often encounter two primary limitations: 1) Redundant or ambiguous textual descriptions: existing methods rely on raw or generated descriptions without filtering, introducing redundant and ambiguous semantic noise; and 2) Underutilization of hierarchical visual features: most approaches align single-layer visual features with auxiliary semantic embeddings, underutilizing hierarchical information. To address these challenges, we propose a task-oriented multimodal FGVC framework that eliminates textual redundancy while enhancing multi-layer alignment between cross-modalities. Specifically, our method comprises two key components: Hierarchical Semantic Purification (HSP) and Multi-layer Cross-Modal Alignment (MCA). The former employs a semantic distillation dictionary to eliminate redundant elements and uses a self-attention mechanism for ranking and semantic refinement. The latter establishes effective cross-modal fusion by integrating multi-layer features with purified text features, effectively combining multi-scale visual representations. Experimental results on 7 public datasets demonstrate that our proposed method outperforms existing counterparts, contributing to advancements in fine-grained visual classification.

Meng-Huan Zhang, Qing Cai, Fan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.