Skip to content
Review Open access

A Review of Multimodal Large Language Models: Fusion Mechanisms and Capability Evolution

Jul 2026 · Computers and artificial intelligence · Vol 3, pp. 167-176 · 0 citations · 32 references

TL;DR

This paper reviews the principal technical paradigms of multimodal fusion, including early fusion, intermediate fusion, late fusion, and hybrid fusion, and compares the structural characteristics and applicable scenarios of different fusion approaches and explores the development of MLLMs from the perspectives of vision-language understanding, multimodal content generation, multimodal interaction and agent-oriented tasks, as well as domain-specific applications.

Abstract

With the continuous advancement of large language models, MLLMs have gradually emerged as a major research focus in the field of artificial intelligence. Although existing studies have reviewed relevant progress from different perspectives, such as multimodal learning, vision-language models, and domain-specific applications, there is still a lack of systematic analysis of the intrinsic relationships among fusion mechanisms, capability formation, and application expansion in MLLMs. This paper reviews the principal technical paradigms of multimodal fusion, including early fusion, intermediate fusion, late fusion, and hybrid fusion, and compares the structural characteristics and applicable scenarios of different fusion approaches. Building on this analysis, the paper further explores the development of MLLMs from the perspectives of vision-language understanding, multimodal content generation, multimodal interaction and agent-oriented tasks, as well as domain-specific applications, thereby outlining the evolutionary trajectory of their core capabilities and broader application trends.

Read PDF

Similar papers

Review Open access Aug 2026

The Versatility of Large Language Models: A Comprehensive Review and Structured Survey of Architectures, Applications, Challenges, and Future Trajectories

This survey reviews the evolution of language models from early statistical approaches to modern Transformer-based architectures and summarizes key developments, including attention mechanisms, scaling laws, alignment techniques, and efficient inference methods.

P. Peykani, V. Charles, Ali Emrouznejad et al. · 0 citations
Conference Open access 2026

Large Language Model Technologies: Progress, Problems and Prospects

Large language models (LLMs) are built on the classic Transformer architecture and have become a core driving force for the rapid development of modern artificial intelligence. This paper presents a systematic review of LLMs, elaborating on their fundamental working principles, mainstream open-source models, effective lightweight optimization methods, retrieval-augmented generation frameworks and key human-value-aligned technologies. Nowadays, LLMs have been widely applied in practice. Typical scenarios include intelligent text generation, professional knowledge-based question answering and automated code generation, delivering remarkable value to both industries and academia. However, their large-scale industrial application is still restricted by multiple challenges. The major issues involve content hallucination, poor model interpretability, excessive computing resource consumption, potential ethical risks and unsatisfactory multimodal integration capability. This paper also forecasts the future development directions of LLMs, such as lightweight deployment on edge devices, safety-focused human value alignment, in-depth cross-modal fusion and customized large models for vertical industries. Additionally, it collects a number of representative cases, which can offer solid references and practical guidance for relevant researchers and engineering practitioners to carry out further studies.

Siyi Fan · 0 citations
Review Open access Jul 2026

Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning

Multimodal video understanding (MVU) has emerged as a fast-growing research frontier, driven by major advances in video-language pre-training and large multimodal models over the past decade. MVU aims to synergistically integrate visual, audio and textual modalities to interpret complex video semantics, supporting widespread downstream tasks including cross-modal retrieval, dense captioning, video question answering, event analysis and intelligent assistance. Despite the rapid proliferation of specialized MVU models, the community still lacks a unified capability-centric framework to systematically clarify the hierarchical competency architecture and evolutionary trajectory of state-of-the-art approaches. To address this issue, this paper presents a structured, comprehensive survey of the latest MVU progress, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning. Along this pipeline, we further systematically synthesize core modality fusion strategies, mainstream benchmark datasets and standardized evaluation protocols. Through a fine-grained analysis of representative published results, we highlight the critical impact of inconsistent evaluation settings, cross-experiment comparability bottlenecks and inherent methodological trade-offs between performance and efficiency. Finally, we identify and dissect three key open challenges: ultra-long video scalability, performance degradation from modality noise and missing data, and factual reliability risks in generative MVU systems. This capability-oriented systematic reference clarifies the methodological evolution logic of MVU, and provides actionable guidance for developing next-generation robust, high-performance multimodal video understanding systems.

Rongyong Zhao, Da Pu, Cuiling Li et al. · 0 citations
#artificial intelligence Review Open access Nov 2026

A comparative review of modern large language model paradigms: GPT-4, BERT, Gemini, and DeepSeek

Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).

Kavish Sanghvi, Aparna S. Sharma, Surbhi Hooda · 0 citations
Review Open access Jul 2026

Optimising Domain-Specific Neuron Activation for Efficient Multimodal Language Understanding In Cloud AI Systems

A unified four-quadrant taxonomy of efficiency strategies is proposed, an integrated future-research agenda built on three converging innovations: adaptive cross-modal attention re-weighting, knowledge-injection pathways, and sparse domain-conditioned neuron gating are outlined, and the cloud-aware evaluation framework that would validate them are outlined.

Olom Ogar Austin, Joshua Abah, Ali Muhammad et al. · 0 citations