Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but dire...
K. M. Le, H. Pham, Danh Thanh Luu et al.· 1 citation
Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or...
H. Pham, P. Tran, Thuan Duc Mai et al.· 1 citation
The GENAI4E team's solution to AI City Challenge 2026 Track 4 builds upon a strong retrieval backbone and progressively integrates heterogeneous vision-language embedding models through score alignment and iterative ensemble fusion, followed by disagreement-aware VLM reranking for ambiguous queries.
Huu-An Vu, Cam Tu Nguyen Thi, Thanh Toan Le Ngo et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.