The GENAI4E team's solution to AI City Challenge 2026 Track 4 builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries.
Abstract
Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural language descriptions. Compared with conventional text-based person retrieval, this task requires fine-grained reasoning over pedestrian appearance, behaviors, object interactions, and scene context, making robust cross-modal matching significantly more challenging. This paper presents the GENAI4E team's solution to AI City Challenge 2026 Track 4. Our framework builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries. On the official Pedestrian Anomaly Behavior (PAB) benchmark, our approach achieves 90.92% mAP, 85.13% Recall@1, 97.72% Recall@5, and 98.68% Recall@10, demonstrating the effectiveness of combining complementary vision-language representations with selective multimodal reasoning for large-scale text-based person anomaly retrieval.
Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or...
H. Pham, P. Tran, Thuan Duc Mai et al.· 1 citation
Text-based person anomaly search requires distinguishing individuals based on fine-grained, context-dependent behaviors rather than mere appearance. Existing methods struggle to capture these context-conditioned actions, frequently relying on isolated skeletal geometry, discarding raw query details during reformulation...
Thanh-Khoi Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen et al.· 0 citations
Text-Based Person Retrieval (TBPR) aims to locate a person in an image database based on a natural language description. While effective in theory, TBPR faces substantial challenges in real-world scenarios due to noisy correspondences—misaligned or weakly related image-text pairs—that significantly degrade retrieval pe...
Zequn Xie, Chu-Xin Wang, Sihang Cai et al.· ACM Transactions on Informat...· 0 citations
An Ambiguity-Aware Semantic Fusion Framework (AAK-LASFNet) for robust text classification and demonstrates that the proposed framework consistently outperforms strong baselines in terms of Accuracy, highlighting its effectiveness in alleviating semantic ambiguity in text classification.
Peijun Xie· Poster Volume 0007 The 2026...· 0 citations
The proposed Cross-Modal Adaptive Token Selection and Alignment Network (CATSANet), a CLIP-based framework tailored for TI-ReID, achieves competitive performance in terms of Rank-k accuracy and mAP, demonstrating the effectiveness of fine-grained alignment and ranking refinement across datasets.
Dongbin Chen, Junjie Li, Hao Xu et al.· Pattern Analysis and Applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.