TIV: A Tri-Modal Text–Image–Video Architecture for Crowd Behavior Classification and Captioning
Understanding crowd behavior in video requires simultaneously reasoning over temporal dynamics, spatial appearance, and semantic context, challenges that single-modality models address only partially. Current vision-language systems suffer from a representation bottleneck originating not from encoder capacities, but fr...