3 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Dual-Prototype Disentanglement Learning for Vision-Language Skeleton Representation
Auxiliary visual and language modalities can improve self-supervised skeleton action representation learning by supplying object, scene, and semantic cues that joint coordinates lack. Existing cross-modal training signals are often defined at the sample level or aligned in a single global space, making them sensitive to noisy external features and prone to suppressing skeleton-specific cues. This paper proposes DVLS, a disentangled vision-language-guided skeleton representation framework built upon a prototype-augmented baseline. DVLS splits each projected modality feature into shared and private subspaces: shared dimensions use cross-modal prototypes to capture transferable semantics, whereas private dimensions use modality-wise prototypes to preserve modality-specific structure. This design reduces the adverse effect of noisy global alignment while retaining external vision-language supervision. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD show consistent gains over a strong global-prototype baseline on four of five skeleton-only linear protocols, including +1.18 points on PKU-MMD XSub and +0.39 points on NTU120 XSub, while matching the baseline on NTU60 XView. Under 1% semi-supervised NTU60 XView, DVLS improves the baseline from 72.04% to 73.35%.
CrossVLS: Cross-Modal Vision-Language Prototypes for Self-Supervised Skeleton Action Representation Learning
Results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines, and CrossVLS is proposed, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference.