Jul 2026· International Journal of Image and Graphics· 0 citations
TL;DR
The research proposes the Fight-Or-Free Optimized Distributed Patch-Wise Attention-Driven Bidirectional Long Short-Term Memory Network (F2DPAB-Net) for visual classification, which outperforms existing methods, thus attaining a maximum of 0.981 Cohen's Kappa Score, 0.96 MCC, and 0.984 NPV.
Abstract
Continuous advancements of deep learning techniques have profoundly influenced Artificial Intelligence (AI) for visual classification through shifting the field from manual feature engineering to autonomous, hierarchical feature learning. On the contrary, the traditional mechanisms for visual classification relied on various challenges, including large data requirements, computational demands, model interpretability issues, and bias concerns, which severely limited accurate classification. Therefore, the research proposes the Fight-Or-Free Optimized Distributed Patch-Wise Attention-Driven Bidirectional Long Short-Term Memory Network (F2DPAB-Net) for visual classification. The Fight-Or-Free Optimization (F2Opt) algorithm significantly tunes the hyperparameters using stochastic behaviors, potentially improving convergence speed and providing a balance between local exploitation as well as global exploration. Integration of patch-wise triplet attention fusion mechanism offers parallel processing, making the model more efficient in learning long-range dependencies that substantially increase the significance while training. On top of that, utilization of multimodality features enables the model to process and understand different modalities that achieve a more comprehensive interpretation of information and improve generalization. Overall, the proposed F2DPAB-Net outperforms existing methods, thus attaining a maximum of 0.981 Cohen’s Kappa Score, 0.96 MCC, and 0.984 NPV using the COCO dataset, respectively.
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that...
Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on...
Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi· 0 citations
Fine-grained visual classification relies on subtle local cues and is highly sensitive to input resolution, yet practical deployment often constrains image size and inference cost. Existing low-resolution recognition methods often depend on additional network structures or complex training modules, limiting deployment...
Chen-Qiang Li· Applied and Computational En...· 0 citations
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
Asynchronous Learning Rate is proposed, which assigns depth-dependent initial learning rates to network modules according to their topological depth, and Smoothed Synchronous Decay is proposed, which coordinates the subsequent decay of heterogeneous parameter groups.
Qiang He, Qiu Zong, Yi-Qi Wang et al.· IEEE Access· 0 citations
ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone, drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language r...
Daniel A. Perkins, J. Squires, Janou Milligan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.