Aug 2026· International Conference on Optoelectronic Information and Computer Engineering (OICE)· Vol 14317, pp. 1431717 - 1431717-13· 0 citations· 40 references
Engineering
TL;DR
LatentAdv, an adversarial camouflage generation framework that integrates ControlNet-based reference style generation and variational autoencoder (VAE) latent space optimization, is proposed, which provides a feasible solution for balancing attack performance and visual naturalness in adversarial object detection.
Adversarial Robustness with Manifold-Oriented Training (ARMOR), a novel defense that realizes the core insights of on-manifold adversarial training (OMAT) in low-data regimes and translates insights from manifold-based training to defend object detectors amidst training data scarcity.
Haoran Wang, Matthew Lau, Alec Helbling et al.· 0 citations
This work proposes an efficient and effective adversarial attack detection scheme leveraging the multi-task perception within a complex vision system, and develops a consistency score metric to measure the inconsistency between vision tasks.
Cong Chen, J. Monteuuis, Jonathan Petit· 0 citations
Adversarial examples generated on convolutional neural network (CNN) surrogates often transfer less effectively to vision transformer (ViT) targets than to CNN targets, creating a cross-architecture bottleneck for transfer-based black-box attacks. Existing input-transformation attacks diversify gradient estimation, yet the evaluated baselines still exhibit substantial CNN-to-ViT transfer gaps, motivating a deformation strategy that varies both control-point geometry and boundary support. To address this limitation, this paper proposes the Enhanced Deformation Attack (EDA), which estimates attack gradients over stochastically transformed views of the current adversarial image. Each view samples either a full control-point grid with movable boundary points or an interior center grid, and the displaced control points are remapped to a reflection-padded canvas before thin-plate spline (TPS) resampling. Gaussian noise or brightness adjustment provides complementary appearance variation. All transformations are confined to gradient estimation, so the final adversarial example remains in the original coordinate system and satisfies the prescribed ℓ∞ perturbation budget. On the ImageNet-Compatible dataset, EDA achieves the highest mean attack success rate (ASR) on ViT targets for all four CNN sources, outperforming the strongest source-specific baseline by 5.8, 11.2, 8.7, and 8.0 percentage points under ResNet-18, Inception-v3, Inception-v4, and Inception-ResNet-v2, respectively. The gains also persist on the full ImageNet validation set and ImageNet-V2 across all evaluated source and target groups. Additional evaluations across four perturbation budgets, targeted settings, five random seeds, defended models, and modern robust models further support the stability and scope of the observed transfer gains. The source code is publicly available at https://github.com/wjc2400136/EDA.
Visual language models (VLMs) have demonstrated outstanding performance in high-value domains such as autonomous driving, unmanned system navigation, and intelligent question-answering; however, the security of their cross-modal alignment mechanisms has not yet been fully verified. Existing visual adversarial patch attacks typically rely on continuous, dense pixel perturbations, which are easily detected and blocked by anomaly detection systems in practical engineering applications. This paper proposes a novel sparse adversarial patch attack algorithm (Sparse Patch Attack, SPA), which successfully misleads the text generation results of VLMs by generating highly dispersed discrete pixel perturbations in non-salient regions of the image. To achieve this, we introduce a differentiable L0-norm approximation and a cross-attention masking mechanism to minimize the number of modified pixels. Furthermore, addressing the characteristics of large-scale model open-ended text generation, we construct a multi-dimensional robustness evaluation framework covering semantic deviation, target achievement rate, and visual concealment. Preliminary experiments on mainstream visual-language models (such as LLaVA and BLIP-2) demonstrate that the SPA algorithm can achieve high success rates in targeted cross-modal attacks with an extremely low pixel modification rate (<1%). This study reveals a novel security vulnerability in visual-language models within complex real-world environments and provides a quantitative evaluation benchmark for future defense mechanisms in multimodal models.
T.-Y. Chen, X.-Y. Hu, J.-F. Wang et al.· Advanced Electromagnetics· 0 citations
AdROD outperforms five baseline defenses and exhibits superior generalizability compared with the evaluated adversarial-training baselines, while maintaining real-time performance for safely stopping the vehicle at a stop sign instrumented with adversarial patches.
Yuting Wu, Dongfang Guo, Xiangzhong Luo et al.· 0 citations
This paper reveals that samples generated by a well-trained generative model are close to clean ones but far from adversarial ones, and proposes Consistency Model-based Adversarial Purification (CMAP), which optimizes vectors within the latent space of a pre-trained consistency model to generate samples for restoring clean data.
Shuhai Zhang, Jiahao Yang, Hui Luo et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.