Skip to content
Open access

TIV: A Tri-Modal Text–Image–Video Architecture for Crowd Behavior Classification and Captioning

2026 · IEEE Access · Vol 14, pp. 139595-139605 · 0 citations · 18 references

Abstract

Understanding crowd behavior in video requires simultaneously reasoning over temporal dynamics, spatial appearance, and semantic context, challenges that single-modality models address only partially. Current vision-language systems suffer from a representation bottleneck originating not from encoder capacities, but from under-exploited correlation spaces during feature fusion. Because multi-stream relational complexity scales combinatorially rather than additively, traditional two-stream designs inherently limit cross-modal synergy. To exploit these high-dimensional relationships, we propose TIV (Text–Image–Video), a fusion-centric tri-modal framework dedicated entirely to optimized joint encoding. Rather than introducing new feature extractors, TIV orchestrates frozen heterogeneous experts by integrating temporal video data via VideoMAE alongside localized spatial cues via a CLIP image encoder, and a text-based semantic guide via a training-only CLIP text encoder, to maximize structural alignment across three distinct data streams. A shared fusion module and a three-way alignment objective make this expanded correlation space learnable, driving both a classification head and a conditional caption decoder, while a modality-dropout curriculum distills the space into the video stream so that the deployed model consumes only raw video. Evaluated on the Crowd-11 benchmark under a strictly scene-disjoint (source-video-grouped) split that eliminates train/test leakage, TIV attains a video-only accuracy of 75.1% (macro-F1 0.750), measured at inference time on the held-out test set with raw video as the sole input, surpassing the strongest video-only baselines VideoMAE (67.7%) and TimeSformer (56.0%) on the identical split. Since that VideoMAE baseline is the same Kinetics-pretrained backbone TIV uses, fine-tuned unimodally, the 7.4-point gain is attributable to the tri-modal training space. We report the full modality breakdown transparently. Only when text and image are additionally supplied at inference does the model reach 93.5% accuracy, a training-time upper bound that reflects the additional semantic information in the caption stream rather than a deployable result. For captioning, we introduce TIVScore, an extension of CLIPScore that incorporates both the image and video modalities into the reference-free similarity computation, providing a more temporally grounded signal than the image-only CLIPScore. TIV obtains a TIVScore of 1.256 (scale $[{0,2.5}]$ ), BLEU-4 of 0.131, CIDEr of 0.583, and ROUGE-L of 0.395, while additionally producing natural-language descriptions for clips it classifies. We shall publicly release code, model weights and the Crowd-11 caption annotations.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.