Size-aware contrastive learning for unbiased video scene graph generation
Video Scene Graph Generation (VidSGG) aims to parse subject-predicate-object triplets from videos, a cornerstone for high-level video understanding. However, existing methods are plagued by severe predicate imbalance: a few frequent predicates (e.g., looking at) dominate the training distribution, leading to heavily bi...