Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining
This work introduces a weakly-supervised vision-language pretraining mechanism that transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without tra...