Skip to content

Cross-Modality Causal-Aware Hierarchical Representation Learning for Domain Generalized Object Detection.

Sep 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP · 0 citations
Medicine

Abstract

Domain Generalized Object Detection (DGOD) addresses the critical challenge of detecting objects across diverse unseen visual domains. Recent advances in vision-language models (VLMs) have shown promising zero-shot generalization capabilities that benefit DGOD. However, existing VLM-based DGOD methods primarily leverage VLMs for data augmentation, overlooking the rich generalizable knowledge it contains. Besides, VLM-based methods developed for other domain generalization (DG) tasks suffer from modality gap and representation bias due to their correlation-driven cross-modal interaction paradigms, which severely limit the performance to DGOD. To bridge this research gap and advance the causal-driven VLM-based DG, we develop a Causal-aware Hierarchical Representation Graph that reformulates the VLM-based DG problem into a hierarchical causal representation learning framework. Our framework incorporates a Generalizable Knowledge Transfer module to refine and inherit transferable scene-object features from the VLM's visual encoder, and a Causal Prototype Learning module that employs general text embeddings as causal interventions to guide the construction of a mediator visual causal prototype space, which inherits the generalization and category relational representation ability from text without representation bias. Furthermore, we introduce a Prototypical Cross-attention Classifier that eliminates modality gap by integrating object features with learned causal prototypes for text-free classification, which also directs visual features to approach the mediator causal space, enabling the learning of causal visual features that are invariant to domain-specific confounders. Our causal-driven framework transfers causal invariance from text to visual modality and provides a vision-friendly perspective for leveraging VLMs to solve vision-centric tasks. Extensive experiments on five benchmarks including Diverse Weather, Corruption, Real-to-Artistic, Cross Camera and Sim-to-Real demonstrate that our method achieves superior generalization performance.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.