Skip to content

Author

Tianzhi Zhu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection

Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as top-down guidance for identifying which weak visual cues are relevant to the target. Recent segmentation systems built on large multimodal models (LMMs) demonstrate this possibility through instruction-conditioned target embeddings that guide mask decoding. However, in this language-to-mask paradigm, the generated target embedding conditions mainly the mask decoder, leaving the dense visual features that must preserve low-contrast boundaries and fine local structure without explicit guidance. We propose Language-Aligned Dense perception for COD (LAD-COD), a framework that aligns top-down semantic target guidance with bottom-up hierarchical visual features. Instead of fully adapting a large generic image encoder, LAD-COD learns a trainable hierarchical visual branch that captures camouflage-sensitive texture, boundary, and contextual information. To align these features with the target embedding, LAD-COD applies Language-Aligned Dual Visual Fusion (LADVF), which extends the embedding beyond sparse prompting to query patch-level language-aligned features and to gate their residual integration with the hierarchical features. This design allows semantic information to guide localization while preserving the fine structural details needed for camouflage segmentation. Experiments on CAMO, COD10K, and NC4K show that LAD-COD obtains the best reported value in all 12 dataset-metric comparisons.

Shangye Song, Tianzhi Zhu, Syed Ariff Syed Hesham et al. · 0 citations
Sep 2026

Understanding Multimodal Learning From Modality Fusion and Alignment Perspectives.

Despite significant advancements in multimodal learning (MML), it has been unexpectedly shown to underperform compared to unimodal approaches in practice, largely due to the modality imbalance problem, ultimately affecting the overall performance of the model. Naturally, most existing methods aim to rebalance optimization speeds across different modalities to avoid performance degeneration caused by modality imbalance. However, in addition to task-oriented modality fusion, we experimentally find that multimodal learning requires explicit modality alignment to stimulate weak modal capabilities so that they can be fully exploited, which is ignored by existing works. Therefore, in this paper, we explore the impact of modality fusion and alignment on multimodal learning from a unified perspective, and develops a dynamic strategy that jointly optimizes both, with particular emphasis on addressing modality imbalance. Concretely, we initially design a soft alignment strategy to impose the positive intervention from the prediction level by integrating modality fusion and alignment into a unified framework. We further extend this strategy to the representation level and hybrid level, enabling compatibility with a wider range of architectures. Subsequently, we design a heuristic strategy to dynamically integrate fusion and alignment. Furthermore, we develop a learning-based strategy using a bi-level optimization framework and theoretically prove the convergence of the learning algorithm to ensure its reliability. These two dynamic integration strategies are incorporated into a unified framework applicable to both supervised and semi-supervised scenarios, further enhancing performance. We conduct a series of experiments to demonstrate the effectiveness of our method on diverse datasets. Extensive experiments demonstrate that our method consistently outperforms state-of-the-art multimodal learning approaches, achieving accuracy improvements of 1.30%, 2.69%, and 0.60% on representative bimodal benchmarks, namely KSounds, CREMA-D, andSarcasm, respectively, as well as gains of 1.35% and 0.65% on trimodal datasets, namely NVGesture and IEMOCAP.

Yang Yang, Fengqiang Wan, Qingjun Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.