Aug 2026· PeerJ Computer Science· 0 citations· 19 references
TL;DR
This work explores different neural network architectures, including fully connected networks, classical convolutional networks, and residual networks, under four types of adversarial attacks constrained by different L p norms, and investigates how adversarial examples affect the internal representations of networks by analyzing the nearest neighbors and class manifold proximity across layers.
Abstract
Deep learning models have achieved remarkable success across various domains, yet they remain vulnerable to adversarial examples, small carefully crafted perturbations of input images that cause models to make incorrect predictions. These adversarial examples are usually indistinguishable from the original input, yet the model classifies them incorrectly, which implies the lack of robustness of trained models. This work explores different neural network architectures, including fully connected networks, classical convolutional networks, and residual networks, under four types of adversarial attacks constrained by different L
p
norms. We evaluated attack success rates across multiple datasets and observed how different models behave when faced with various adversarial examples. All attacks are remarkably effective across all models and lead to misclassification almost every time. Next, we investigate how adversarial examples affect the internal representations of networks by analyzing the nearest neighbors and class manifold proximity across layers. Our results show that misclassification often occurs in the last couple of layers of the models, with variations depending on the dataset and the model used. In order to use a large model such as Residual Network 18 (ResNet-18), we apply principal component analysis to reduce unnecessary dimensions and to lower time complexity. We also analyzed how this reduction affects the results. This work highlights the importance of understanding not only if a model fails under a given attack but also how and where these failures occur within the network architecture.
A method to analyze ANNs designed for image classification from an adversarial robustness perspective and implemented an ablation and fine-tuning strategy that successfully boosted the robustness of the ANNs against a variant of the Auto-PGD attack under different threat models.
Among the evaluated models, CNNs exhibit the highest baseline robustness, whereas DNNs and RNNs rely more heavily on defense mechanisms to maintain performance, whereas DNNs and RNNs rely more heavily on defense mechanisms to maintain performance.
Surekha M., A. K. Sagar, Vineeta Khemchandani· International Journal of Int...· 0 citations
Deep neural networks achieve impressive performance in image and speech recognition, yet they are sensitive to small input perturbations known as adversarial examples that can cause critical misclassifications. This vulnerability motivates classification systems that are inherently stable and robust. In this work, we f...
Israe El Ghizi, Abdellah Ait Omar, Khawla Hamouichou et al.· EPJ Web of Conferences· 0 citations
The research methodology involved a systematic literature review using the Scopus database, adhering to Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines, and focusing on recent advancements in attack and defence techniques.
The study concluded that adversarial resilience is largely determined by the interaction between model architecture and defense strategy, highlighting the need for architecture-specific defense selection when developing secure medical image classification systems.
Y. Heryadi, I. Sonata, Bambang Krismono Triwijoyo· Matrik· 0 citations
That such a regularizer exists is the main finding: the methods that dispense with the inner search all obtain their local geometry by differentiating with respect to the input, and it is shown this is not necessary.
T. C. Johnson, Donsub Rim· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.