Skip to content

Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning

Sep 2026 · 0 citations · 68 references
Computer Science

TL;DR

PolyMem, an exemplar-free approach that implicitly models rich high-order statistics of the feature distribution to enhance cross-domain robustness, is introduced that effectively alleviates the performance discrepancy while improving the model's performance across domains.

Abstract

3D perception plays a crucial role in real-world applications such as autonomous driving, robotics, and AR/VR. In practical scenarios, 3D perception models need to continually adapt to newly emerging 3D object categories, making class-incremental learning (CIL) particularly important. However, unlike 2D images, 3D point clouds are inherently heterogeneous: objects from the same class may not only come from the clean CAD domain, but also from RGB-D camera scans of varying quality, video reconstructions, or even corrupted observations. We discover that such heterogeneity introduces a new challenge beyond catastrophic forgetting: the degree of performance degradation can vary substantially across domains, a phenomenon we term performance discrepancy. To investigate this problem, we establish the Domain3D-CIL training and evaluation protocol, which contains point cloud categories from heterogeneous domains. We further adapt a wide range of mainstream CIL methods to the 3D modality. The results demonstrate that this performance discrepancy consistently appears across these baselines. To mitigate this issue, we introduce PolyMem, an exemplar-free approach that implicitly models rich high-order statistics of the feature distribution to enhance cross-domain robustness. Experiments demonstrate that our method effectively alleviates the performance discrepancy while improving the model's performance across domains. Code will be made publicly available upon acceptance.

View source

Similar papers

Preprint Sep 2026

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, w...

B. Duisterhof, Kai-Feng Zhang, Adam Hung et al. · 0 citations
Preprint Sep 2026

OTT3R: Multi-View 3D Reconstruction and Fast Dataset Generation at 1% Compute

Feed-forward 3D reconstruction models have achieved impressive performance by scaling model and dataset size, but their cost excludes most research groups and precludes edge deployment. Additionally, generating 3D supervision without sensors still relies on slow, unreliable Structure-from-Motion, as the community lacks...

Brandon Leblanc, Charalambos Poullis · 0 citations
Preprint Sep 2026

DIFTA-3D: Depth-Consistent Instance-Level Feature Transfer and Adaptation of DINOv3 for 3D Detection

RGB-D 3D instance detectors benefit from visual semantics, but the task-specific Faster R-CNN/ResNet branch used by IIFNet3D couples feature extraction to a separately trained 2D detector and its image-domain labels. Replacing that branch with a frozen vision foundation model removes this task-specific dependency, but...

Lin-Man Wang, Zi-Fei Zhang, Chun-Ran Zheng et al. · 0 citations
Preprint Aug 2026

SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

SpatialCrafter is presented, a novel two-stage framework that addresses explorable image-to-scene generation issues by introducing a global 3D proxy for high-fidelity image-to-scene generation and appearance refinement and introduces Parallel Geometry Injection and Proxy-Aware Corruption training strategies.

Chuan Fang, Lingteng Qiu, Yixun Liang et al. · 1 citation
Preprint Aug 2026

MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection

Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models...

M. Kotb, Johannes Meier, Christoph Reich et al. · 0 citations
2026

PointLIBERO: Unlocking Spatial Awareness in VLAs With a Novel 3-D Dataset and a Lightweight Framework

Vision-Language-Action models such as OpenVLA and DexVLA have demonstrated impressive generalization by leveraging large-scale 2D robotic datasets. However, their reliance on 2D RGB imagery significantly limits their 3D spatial reasoning, leading to spatial naivety in depth-sensitive tasks. Bridging this gap is hindere...

Meng Li, Lijiang Chen, Qi Zhao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.