Aug 2026· Physica Scripta· Vol 101, pp. 326002· 0 citations· 58 references
Physics
TL;DR
This work introduces saccadic lightweight efficient attention (SLEA), a framework purpose-built for 2D medical image classification and inspired by selective visual processing in the human visual system, intended to reduce dependence on fixed image locations and recurring spatial artefacts.
Abstract
Deep learning in medical imaging faces a persistent tension between preserving localized diagnostic information and maintaining computational efficiency. We introduce saccadic lightweight efficient attention (SLEA), a framework purpose-built for 2D medical image classification and inspired by selective visual processing in the human visual system. Rather than dividing an image into fixed patches or processing it at once, SLEA independently samples multiple localized crops from uniformly distributed spatial coordinates, with sampled locations re-drawn at every training epoch to provide spatially diverse exposure without requiring a learned fixation policy. Each glimpse is encoded through a shared ImageNet-pretrained MobileNetV2 backbone into a compact feature token; the resulting unordered token set is integrated through multi-head self-attention followed by mean pooling, capturing content-level relationships among spatially distributed regions within a minimal parameter budget. This stochastic sampling is intended to reduce dependence on fixed image locations and recurring spatial artefacts, although this regularization effect is an empirical hypothesis rather than a formally established mechanism. Experiments on ChestX-ray14, ISIC2019, and APTOS2019 show strong performance under the evaluated protocols while maintaining a compact model size of approximately 2.45 million parameters. However, the ChestX-ray14 result is substantially higher than the commonly reported benchmark range and should be interpreted cautiously; independent replication under closely matched partitions, preprocessing, optimization, and inference protocols remains necessary. Results from the COVID-19 Radiography Database are reported separately, since publicly assembled COVID-19 datasets may contain source-, acquisition-, and processing-related differences that create artificial class separability; this result is therefore excluded from the principal performance claims and is not interpreted as evidence of clinical diagnostic capability or general superiority. A multi-granularity interpretability scheme combining glimpse-level Grad-CAM, top-glimpse, and multi-glimpse visualizations provides qualitative indications of the regions associated with the model’s predictions, but these are not interpreted as exact lesion-localization maps or definitive evidence of clinically faithful reasoning.
LGF-Net is proposed, a unified framework that jointly models spatial semantics and adaptive spectral cues for deepfake detection and achieves competitive intra-dataset and cross-dataset performance compared with several state-of-the-art deepfake detection methods.
Experimental results demonstrate that adaptability is not solely determined by model size, but rather by how effectively parameter plasticity is regulated in dynamic environments.
Xiao-Rong Zeng, Weiqiang Chen, Peng Shi et al.· 0 citations
Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations have already been la...
Zhe-Han Kan, Xing-Hua Jiang, Yu-Bo Zhu et al.· 0 citations
Implicit Neural Representations (INRs) provide a flexible and resolution-independent formulation for continuous signal representation. Despite their strong representation ability, standard INRs usually predict each queried coordinate independently, making it difficult to explicitly exploit the local coherence widely ob...
Chen Qing, Wen-Xin Zhang, Hao-Yu Wang et al.· Applied Sciences· 0 citations
Recent advances in single image super-resolution (SISR) have leveraged convolutional neural networks (CNNs) and vision transformers to model pixel-level statistics, often relying on increasingly complex architectures to capture spatial correlations. However, these approaches generally overlook a fundamental distinction...
Qi-Bin Zhang, Li-Cheng Liu, Ting-Yun Liu et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.