Skip to content
Open access

CrossVLS: Cross-Modal Vision-Language Prototypes for Self-Supervised Skeleton Action Representation Learning

Jul 2026 · Electronics · 0 citations · 36 references

TL;DR

Results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines, and CrossVLS is proposed, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference.

Abstract

Artificial intelligence (AI)-driven positioning and tracking systems combine geometric trajectories with behavior understanding in smart-city, healthcare, and autonomous environments. Skeleton sequences provide compact, privacy-preserving motion geometry, but labeled data are costly, and coordinate-only self-supervision cannot recover object, scene, or interaction cues. Vision-language transfer can supply these cues, but instance-level targets remain sensitive to noisy crops, incomplete descriptions, and ambiguous actions. We propose CrossVLS, which is a cross-modal vision-language-guided framework that transfers semantic knowledge from red, green, and blue (RGB) frames and generated language descriptions to a skeleton encoder during pretraining while retaining skeleton-only inference. CrossVLS replaces noisy instance-level transfer with a shared prototype space: skeleton, RGB, and language features are softly assigned to a common prototype bank through balanced optimal transport, and the resulting assignments define semantic soft targets for contrastive learning. A full-batch progressive training schedule gradually increases cross-modal guidance without splitting the physical batch, preserving the support set used to construct semantic targets. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD demonstrate strong performance under linear and semi-supervised evaluation using only the pretrained skeleton encoder at inference. These results show that prototype-mediated vision-language transfer can improve skeleton representations for the behavior-interpretation stage of positioning and tracking pipelines.

Read PDF

Similar papers

Aug 2026

Self-supervised skeleton action recognition based on graph prototype learning

This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.

Zhijie Xu, Hong-Wei Chen, Xia Li · 0 citations
Open access Aug 2026

SIRModel: Learning Spatial Intermediate Representation to Parameter-Efficiently Fine-Tune a Vision Language Model for Manipulation

Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which rely on expensive human annotations and training cost. This article presents a spatial Gaussian-guided hierarchical framework that uses ordered 3D Gaussian interaction regions as an explicit planning interface between vision–language reasoning and action generation. The proposed framework enables efficient adaptation of a pretrained vision–language model for robotic manipulation tasks. First, an automatic geometric enhancement pipeline converts raw robot demonstration videos into near-, mid-, and late-stage Gaussian supervision through foreground extraction, metric depth estimation, stable camera aggregation, end-effector localization, 3D lifting, and temporal grouping, without requiring manual 3D interaction annotation. The generated Gaussian representations provide structured spatial guidance, where their covariance characterizes interaction-region extent and variability rather than fully calibrated physical uncertainty. Second, a shared vision–language backbone predicts structured subtasks and Gaussian interaction regions, while a conditional diffusion executor generates future action chunks under these semantic and geometric conditions. A trajectory-to-Gaussian likelihood objective explicitly encourages consistency between generated motions and the predicted spatial interaction plan. Experiments on a mixed real-robot dataset derived from LHManip and RH20T show that our method improves trajectory tracking success from 55.7% to 70.8% over a same-backbone direct VLA baseline. Closed-loop simulation evaluation on LIBERO with 80% backbone parameter frozen achieves 85.3% average task success, demonstrating the effectiveness of explicit 3D interaction representations for spatial reasoning and long-horizon manipulation.

Li Lin, Ming-Hao Shi, Teng-Long Wang · 0 citations
Preprint Aug 2026

Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition

This proposed TDSM-MM has been ablated via extensive experiments and achieved the best inductive accuracy on three of four NTU-60/120 splits and surpasses the transductive state-of-the-art on NTU-120 96/24, suggesting that diffusion-based methods can be a promising direction for zero-shot learning.

Zehao Bao, Shu-Jun Guo, Bruce X. B. Yu · 0 citations
Jul 2026

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

An object-centric 3D representation alignment framework built upon $\pi_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training, which enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time.

Zong-He Liu, Shan Jie, Xiao-Quan Sun et al. · 1 citation
Preprint Aug 2026

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.

Zijian Zhang, Yu-Qing Jiang, Weitao Zhou et al. · 0 citations
Preprint Aug 2026

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

FineX is introduced, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology and raises mean class accuracy on Gym99, Gym288, and Diving48 without textual supervision or large-scale vision-language pre-training.

Imtiaz ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.