This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions.
Abstract
We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target actions and pretraining using large-scale action scenery datasets. Specifically, our approach, termed Skeleton-Language feature Pooling Switching, introduces a weakly-supervised vision-language pretraining mechanism. This mechanism transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions. Furthermore, we propose Scene-Mixed Discriminative Contrastive Learning to distinguish actions at the instance level within the combined scene through the MIL framework. Our experiments on four public spatio-temporal action localization and classification datasets demonstrate that the proposed method effectively addresses annotation limitations.
A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.
Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al.· IEEE Access· 0 citations
Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which lim...
Hongsong Wang, Peng-Cheng Yan, Yang Zhang et al.· 0 citations
A novel Mask-aware Action Spatiotemporal Quantization framework that decouples the conflicting tasks of spatial feature inference and temporal smoothing, and establishes a comprehensive and substantial leading advantage in the Mean over Frames accuracy.
Xingchen Qin, Lin-Xiang Peng, You-Bao Ye et al.· 0 citations
Unsupervised action segmentation aims to discover latent action categories and their temporal organization without action annotations. Optimal transport-based methods provide structured frame-to-action assignments, however, their pseudo-label quality is fundamentally conditioned on the representation space used to cons...
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hong-Wei Chen, Xia Li· International Journal of Mac...· 0 citations
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-l...
Jia-Xin Zhang, Hao-Ran Gao, Xi-Zhan Gao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.