Skip to content
Review Open access

A Survey of Multi-Model Collaboration in Video Understanding

Sep 2026 · International Conference on Data Technologies and Applications · 0 citations · 114 references

Abstract

The rapid development of multimodal foundation models has shifted video understanding from perception-centered recognition toward more general semantic interpretation, reasoning, and decision-making over dynamic visual content. As video understanding tasks increasingly require fine-grained perception, long-range temporal modeling, multimodal grounding, and adaptive reasoning, collaboration among heterogeneous functional units, including specialized models, modules, agents, memory systems, and external tools, has emerged as an important system-level paradigm. However, existing surveys mainly organize video understanding methods by architectures, learning strategies, or task categories, leaving the collaborative structure of modern systems insufficiently examined. This survey provides a structured narrative review of multi-model collaboration in video understanding, which we formulate as collaborative video understanding. We introduce a unified analytical framework that characterizes collaborative systems through functional units, inter-unit communication mechanisms, and collaborative state representations, and organize existing methods according to their coordination dynamics into static collaboration and dynamic collaboration, with the latter further distinguished into controller-based and agent-based collaboration. We further review representative benchmarks, evaluation metrics, and empirical analysis, showing that current evaluation protocols mainly capture task-level performance but provide limited insight into collaborative organization, memory use, adaptive execution, and system-level collaborative capability. Finally, we discuss key challenges and future directions, including adaptive task decomposition, semantically aligned inter-unit communication, persistent shared memory, uncertainty-aware error containment, evidence-grounded reasoning, and collaboration-centric evaluation. By reinterpreting video understanding from a collaborative systems perspective, this survey aims to provide a structured foundation for developing more adaptive, reliable, and scalable video understanding systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.