Skip to content

F2DPAB-Net: Fight-Or-Free Optimized Distributed Patch-Wise Attention-Driven Deep Learning Network for Visual Classification

Jul 2026 · International Journal of Image and Graphics · 0 citations

TL;DR

The research proposes the Fight-Or-Free Optimized Distributed Patch-Wise Attention-Driven Bidirectional Long Short-Term Memory Network (F2DPAB-Net) for visual classification, which outperforms existing methods, thus attaining a maximum of 0.981 Cohen's Kappa Score, 0.96 MCC, and 0.984 NPV.

Abstract

Continuous advancements of deep learning techniques have profoundly influenced Artificial Intelligence (AI) for visual classification through shifting the field from manual feature engineering to autonomous, hierarchical feature learning. On the contrary, the traditional mechanisms for visual classification relied on various challenges, including large data requirements, computational demands, model interpretability issues, and bias concerns, which severely limited accurate classification. Therefore, the research proposes the Fight-Or-Free Optimized Distributed Patch-Wise Attention-Driven Bidirectional Long Short-Term Memory Network (F2DPAB-Net) for visual classification. The Fight-Or-Free Optimization (F2Opt) algorithm significantly tunes the hyperparameters using stochastic behaviors, potentially improving convergence speed and providing a balance between local exploitation as well as global exploration. Integration of patch-wise triplet attention fusion mechanism offers parallel processing, making the model more efficient in learning long-range dependencies that substantially increase the significance while training. On top of that, utilization of multimodality features enables the model to process and understand different modalities that achieve a more comprehensive interpretation of information and improve generalization. Overall, the proposed F2DPAB-Net outperforms existing methods, thus attaining a maximum of 0.981 Cohen’s Kappa Score, 0.96 MCC, and 0.984 NPV using the COCO dataset, respectively.

View source

Similar papers

Preprint Sep 2026

LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba

While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that...

Li Mu, Shuai Chen, Wen Zheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models

Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on...

Konstantinos D. Polyzos, Eleni Oikonomou, Tara Javidi · 0 citations
Open access Sep 2026

Cross-Resolution Knowledge Distillation for Low-Compute Fine-Grained Bird Classification

Fine-grained visual classification relies on subtle local cues and is highly sensitive to input resolution, yet practical deployment often constrains image size and inference cost. Existing low-resolution recognition methods often depend on additional network structures or complex training modules, limiting deployment...

Chen-Qiang Li · 0 citations
Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations
Open access 2026

Hierarchy-Aligned Learning Rates for Vision Networks

Asynchronous Learning Rate is proposed, which assigns depth-dependent initial learning rates to network modules according to their topological depth, and Smoothed Synchronous Decay is proposed, which coordinates the subsequent decay of heterogeneous parameter groups.

Qiang He, Qiu Zong, Yi-Qi Wang et al. · 0 citations
Preprint Aug 2026

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone, drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language r...

Daniel A. Perkins, J. Squires, Janou Milligan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.