Jul 2026· IEEE International Conference on Engineering of Complex Computer Systems· pp. 1-6· 0 citations· 39 references
Abstract
Auxiliary visual and language modalities can improve self-supervised skeleton action representation learning by supplying object, scene, and semantic cues that joint coordinates lack. Existing cross-modal training signals are often defined at the sample level or aligned in a single global space, making them sensitive to noisy external features and prone to suppressing skeleton-specific cues. This paper proposes DVLS, a disentangled vision-language-guided skeleton representation framework built upon a prototype-augmented baseline. DVLS splits each projected modality feature into shared and private subspaces: shared dimensions use cross-modal prototypes to capture transferable semantics, whereas private dimensions use modality-wise prototypes to preserve modality-specific structure. This design reduces the adverse effect of noisy global alignment while retaining external vision-language supervision. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD show consistent gains over a strong global-prototype baseline on four of five skeleton-only linear protocols, including +1.18 points on PKU-MMD XSub and +0.39 points on NTU120 XSub, while matching the baseline on NTU60 XView. Under 1% semi-supervised NTU60 XView, DVLS improves the baseline from 72.04% to 73.35%.
Results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines, and CrossVLS is proposed, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference.
Kenan Ye, Shengjie Zhao, Shuang Liang· Electronics· 0 citations
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations
This proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.
Zehao Bao, Shu-Jun Guo, Bruce X. B. Yu· 0 citations
FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.
Imtiaz ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al.· 0 citations
Experiments on egocentric video benchmarks show LogFA significantly improves model generalization to unseen environments while maintaining low computational and data collection costs.
A Low-Complexity Cross-Modal Alignment via Projection (LCAP) network is proposed, which introduces Projective Token Compression (PTC), which leverages Mish activation and adaptive average pooling to reduce feature redundancy while enhancing discriminative information, and Positional Spatial Enhancement (PSE), which explicitly injects positional cues into the compressed representations and strengthens spatial structure.
Yu-Chen Sha, Lingli Wan, Ge Yang et al.· The Visual Computer· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.