A Survey of Multi-Model Collaboration in Video Understanding
Abstract
The rapid development of multimodal foundation models has shifted video understanding from perception-centered recognition toward more general semantic interpretation, reasoning, and decision-making over dynamic visual content. As video understanding tasks increasingly require fine-grained perception, long-range temporal modeling, multimodal grounding, and adaptive reasoning, collaboration among heterogeneous functional units, including specialized models, modules, agents, memory systems, and external tools, has emerged as an important system-level paradigm. However, existing surveys mainly organize video understanding methods by architectures, learning strategies, or task categories, leaving the collaborative structure of modern systems insufficiently examined. This survey provides a structured narrative review of multi-model collaboration in video understanding, which we formulate as collaborative video understanding. We introduce a unified analytical framework that characterizes collaborative systems through functional units, inter-unit communication mechanisms, and collaborative state representations, and organize existing methods according to their coordination dynamics into static collaboration and dynamic collaboration, with the latter further distinguished into controller-based and agent-based collaboration. We further review representative benchmarks, evaluation metrics, and empirical analysis, showing that current evaluation protocols mainly capture task-level performance but provide limited insight into collaborative organization, memory use, adaptive execution, and system-level collaborative capability. Finally, we discuss key challenges and future directions, including adaptive task decomposition, semantically aligned inter-unit communication, persistent shared memory, uncertainty-aware error containment, evidence-grounded reasoning, and collaboration-centric evaluation. By reinterpreting video understanding from a collaborative systems perspective, this survey aims to provide a structured foundation for developing more adaptive, reliable, and scalable video understanding systems.