Skip to content
Preprint

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

Aug 2026 · 0 citations · 81 references
Computer Science

TL;DR

This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions.

Abstract

We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target actions and pretraining using large-scale action scenery datasets. Specifically, our approach, termed Skeleton-Language feature Pooling Switching, introduces a weakly-supervised vision-language pretraining mechanism. This mechanism transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions. Furthermore, we propose Scene-Mixed Discriminative Contrastive Learning to distinguish actions at the instance level within the combined scene through the MIL framework. Our experiments on four public spatio-temporal action localization and classification datasets demonstrate that the proposed method effectively addresses annotation limitations.

View source

Similar papers

Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Preprint Aug 2026

Zero-Shot Skeleton-Based Action Anticipation

Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which lim...

Hongsong Wang, Peng-Cheng Yan, Yang Zhang et al. · 0 citations
Preprint Aug 2026

MASQ: Mask-Aware Spatiotemporal Quantization for Unsupervised Skeleton Action Segmentation

A novel Mask-aware Action Spatiotemporal Quantization framework that decouples the conflicting tasks of spatial feature inference and temporal smoothing, and establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.

Xingchen Qin, Lin-Xiang Peng, You-Bao Ye et al. · 0 citations
Preprint Aug 2026

See the Change, Keep the Flow: Unsupervised Action Segmentation via Spectral-Temporal Representation Learning

Unsupervised action segmentation aims to discover latent action categories and their temporal organization without action annotations. Optimal transport-based methods provide structured frame-to-action assignments, however, their pseudo-label quality is fundamentally conditioned on the representation space used to cons...

Yun Li, Jun Xiao, Cong Zhang et al. · 0 citations
Aug 2026

Self-supervised skeleton action recognition based on graph prototype learning

This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.

Zhijie Xu, Hong-Wei Chen, Xia Li · 0 citations
Preprint Sep 2026

Few-Shot Video Recognition via Hierarchical Metric Learning

Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-l...

Jia-Xin Zhang, Hao-Ran Gao, Xi-Zhan Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.