Skip to content

Author

H. Aljuhani

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Multimodal Document Classification Across Domains and Methods: A Systematic Review

Multimodal Document Classification (MDC) is a significant area of research, enabling the integration of textual, visual, and structural features to enhance document understanding across diverse application domains. While substantial progress has been made in developing methods that leverage deep learning, graph-based representations, and cross-modal fusion techniques, the field remains fragmented accross datasets, evaluation protocols, and methodological frameworks. This systematic review synthesizes existing literature on MDC, focusing on methods, architectures, cross-domain adaptability, evaluation metrics, and current challenges. Moreover, following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, the current work surveyed peer-reviewed studies published between 2021 and 2025, identifying trends in feature extraction, fusion strategies, and benchmark datasets. The analysis conducted highlights three key insights: i) the shift from handcrafted features to transformer-based and multimodal pre-trained models; ii) the growing importance of domain-specific adaptations in legal, healthcare, and scientific documents; and iii) persistent challenges related to scalability, interpretability, and generalizability across domains. Overall, this review provides a comprehensive resource for researchers and practitioners, aiming to consolidate knowledge and guide future advancements in MDC.

H. Aljuhani, M. Dahab, Yousef Alsenani · 0 citations