The Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2, achieves state-of-the-art performance and ablation studies confirm the importance of teacher-student semantic compatibility.
Abstract
Vision foundation models (VFMs) offer strong generalization capabilities for domain-adaptive object detection (DAOD). However, existing VFM-based methods overlook the spatial-scale discrepancy between teacher and student feature maps, resulting in semantic incompatibility that weakens both feature alignment and pseudo-label learning. Moreover, domain shift can cause source-trained VFM teachers to miss target-domain objects, limiting the quality of their pseudo-labels. To address these issues, we propose the Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2. SLE Adapter injects pretrained local-texture priors into DINOv2 to improve cross-domain recognition and reformulates its features into dense representations that are spatially and semantically compatible with the student detector. SLE-T transfers the resulting teacher knowledge through either pseudo-label learning or feature alignment. We instantiate SLE-T with DINOv2-B and DINOv2-L (the ViT-B and ViT-L variants) and compare them with the larger DINOv2-G teacher. Extensive experiments on three DAOD benchmarks demonstrate that our method achieves state-of-the-art performance, and ablation studies confirm the importance of teacher-student semantic compatibility. Notably, SLE-T with DINOv2-B produces competitive or superior pseudo-labels using approximately one-quarter of the training time of DINOv2-G and substantially less GPU memory, demonstrating efficient VFM knowledge transfer under limited computational resources.
Source-free domain adaptation (SFDA) aims to adapt pre-trained source models to new target domains without requiring access to any source domain data, thereby addressing privacy and efficiency concerns. Existing SFDA methods for object detection primarily follow a teacher–student self-training paradigm; however, their performance is often limited by noisy pseudo-labels. To address this issue, this paper proposes an SFDA object detection framework for the YOLO family of single-stage detectors. First, a weak–strong pseudo-label consistency filtering strategy is designed to remove unreliable pseudo-labels by exploiting the prediction consistency across different augmented views. Second, a multiscale object-level contrastive learning mechanism is introduced to extract object-level features at multiple feature scales, thereby enhancing the consistency and discriminability of object representations across different views and scales through supervised contrastive constraints. Experimental results show that the proposed method consistently outperforms the baseline on multiple cross-domain detection tasks, demonstrating its effectiveness and good generalization ability under the source-free setting.
This work proposes AMCA, an unbiased SGG framework integrating adaptive multi-prototype learning with cross-modal alignment, which achieves consistently competitive performance across multiple SGG tasks, with particularly strong improvements on the unbiased mR@K metric.
Jinhao Fan, Yuanhao Xi, Chuanping Hu et al.· Journal of King Saud Univers...· 0 citations
Unsupervised Domain Adaptation (UDA) for object detection remains challenging under adverse weather due to significant distribution shifts. While recent Vision Foundation Model (VFM) based methods show promise, they often encounter limitations in extreme domain gaps and pseudo-label noise. This paper proposes two enhancements to the DINO Teacher framework: (1) a multi-round self-training strategy to refine the labeling model progressively, and (2) a depth-guided spatial modulation mechanism using geometric priors from a DINOv3based depth estimator. By modulating the input space, the student model is encouraged to emphasize spatial cues that are less sensitive to visibility degradation in foggy environments. Experiments on Foggy Cityscapes demonstrate that our approach reaches 56.5% mAP with a VGG-16 backbone and 59.8% mAP with ResNet-50. These results demonstrate competitive performance compared with the DINO Teacher baseline and recent Vision Language Model (VLM) based methods, particularly for tail categories such as bus and train.
T. Doan, D. C. Bui, Khanh-Duy Nguyen et al.· International Conference on...· 0 citations
Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using foundation models and then trains a 3D-OVD model based on these discovered objects. Although effective, this pipeline often suffers from inaccurate localization and mismatched classification during the discovery stage, which subsequently limits the performance of the model training stage. To address these limitations, we advocate for improving both the reliability of novel object discovery and the robustness of model training, and propose an innovative framework. Specifically, for reliable discovery, our co-distillation strategy distills high-quality novel objects by applying Hungarian matching over a comprehensive score that incorporates geometric consistency, structural objectness, and semantic certainty. To enhance robust model training, we further propose a dual-guidance learning scheme, incorporating a scene-awareness-guided uncertainty regularization for the regression head and an LLM-guided hierarchical alignment for the classification head, effectively mitigating the negative effects of imprecise 3D bounding boxes and semantic ambiguity. Extensive experiments on SUN RGB-D and ScanNetV2 demonstrate that our method achieves significant performance gains over state-of-the-art approaches. Code is available at https://github.com/shangboyuan/Co-3DGT
Shangbo Yuan, Jie Xu, Xiaofeng Zhu et al.· 0 citations
This work proposes PointPDF V2, a unified framework that integrates open-set recognition (OSS) and incremental learning (IL) into a cohesive pipeline and introduces a more challenging continual OWSS protocol in 3D, where models must simultaneously preserve the known-class performance, acquire new knowledge, and still identify the remaining unknowns across sequential updates.
Jinfeng Xu, Xianzhi Li, Yixue Hao et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.