Jul 2026· International journal of computer information systems and industrial management applications· 0 citations
TL;DR
A unified four-quadrant taxonomy of efficiency strategies is proposed, an integrated future-research agenda built on three converging innovations: adaptive cross-modal attention re-weighting, knowledge-injection pathways, and sparse domain-conditioned neuron gating are outlined, and the cloud-aware evaluation framework that would validate them are outlined.
Abstract
Multimodal large language models (MLLMs) have made it possible for artificial-intelligence systems to reason jointly across vision and language, supporting tasks ranging from image captioning and visual question answering to clinical decision support and autonomous perception. As MLLM scale grows, however, deploying these models in cost- and energy-bounded cloud environments has become a defining engineering challenge. This mini-review consolidates recent literature at the intersection of four research strands: (i) multimodal architecture design and fusion strategies, (ii) neural-activation patterns and mechanistic specialization, (iii) selective and conditional computation including mixture-of-experts, and (iv) cloud-deployment optimisation. We propose a unified four-quadrant taxonomy of efficiency strategies, trace the field's evolution through a decade-scale timeline, and synthesise sixteen primary studies in cross-cutting comparison tables. Particular attention is given to the under-explored intersection of domain adaptation and efficiency, where evidence is converging that neuron-level domain awareness can simultaneously reduce inference cost and improve interpretability. We identify five persistent limitations of current methods and five concrete research gaps that follow from them. The review closes with an integrated future-research agenda built on three converging innovations: adaptive cross-modal attention re-weighting, knowledge-injection pathways, and sparse domain-conditioned neuron gating, and outlines the cloud-aware evaluation framework that would validate them. The article is intended as a single-source reference for researchers and practitioners designing efficient, trustworthy multimodal AI for cloud deployment
Experiments show that MKB achieves competitive scientific understanding across biological and molecular benchmarks, produces high-fidelity native outputs for weather forecasting, biological generation, and medical-image segmentation, and largely retains the general capabilities of its Qwen3-VL backbone.
Hesen Chen, Xinyue Su, Xiaomeng Yang et al.· 0 citations
Neural Architecture Search (NAS) has emerged recently as a powerful paradigm for automating deep neural network design. However, most existing NAS methods are optimised for a single domain, limiting their generalisation to diverse application areas such as computer vision, natural language processing, healthcare, speech recognition, and edge intelligence. This paper proposes a Domain-Adaptive Neural Architecture Search (DA-NAS) framework that learns domain-aware architectural patterns while it maintains a shared search space and optimisation strategy. DA-NAS combines domain embeddings, multi-objective optimisation, and resource-awareness to generate architectures that adapt to heterogeneous data characteristics and deployment constraints. Extensive experiments across multiple domains demonstrate that the proposed approach reduces search cost and improves cross-domain transferability, consistently outperforming domain-specific handcrafted models and conventional NAS baselines.
A. Sindhu Devi, L. Godlin Atlas· International journal of com...· 0 citations
An overview of Vision LLM architectures, their applications and the challenges they face and case study of how building of AI Models through visionLLM may help IndoAI AI camera system are provided.
Rohit Yadav· Journal of Artificial Intell...· 0 citations
A decision-oriented framework that translates theoretical insights into practical guidelines is introduced, establishing a structured foundation for broader communities to train and deploy human-centered AI systems sustainably and efficiently.
C. Villarreal, J. Luzuriaga, Emilio Quinga et al.· AHFE International· 0 citations
Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).
Kavish Sanghvi, Aparna S. Sharma, Surbhi Hooda· Computer Science and Informa...· 0 citations
A definitive taxonomy of the medical VLM landscape is provided, tracing the evolution from early Contrastive Alignment and Generative MLLMs to the cutting-edge frontiers of Dense Pixel-Grounding, Sparse Mixture-of-Experts (MoE), and Reasoning-Incentivized (RL) architectures.
Taha Razzaq, Murtaza Taj, Asim Iqbal· Journal of Biomedical Inform...· 0 citations