Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but have been demonstrated to be vulnerable to corpus poisoning. Existing poisoning attacks against RAG largely focus on single-point explicit injection, where the malicious payload is fully encapsulated within a single document. Consequently, recent mitigation mechanisms have evolved to identify and diminish these threats effectively. In this paper, we first verify that existing mitigation mechanisms are insufficient for a new class of threats: indirect logic induction. Motivated by this observation, we introduce InceptionRAG, a stealthy attack mechanism that subverts the standard attack paradigm. Instead of injecting explicit malicious payloads, InceptionRAG fragments it into a chain of dormant passages. These passages appear harmless and can bypass existing mitigation mechanisms when examined separately. However, when retrieved together, they trigger LLMs to self-deduce target misinformation via multi-hop reasoning. To further improve the applicability of InceptionRAG in black-box settings, we propose zeroth-order suffix optimization (ZOSO) to automate the generation of authoritative suffixes. Extensive evaluations across three datasets and three LLMs demonstrate that InceptionRAG achieves an attack success rate exceeding 80% even under rigorous adversarial constraints. In particular, InceptionRAG shows superior evasion capabilities, effectively bypassing established defenses that mitigate traditional single-document injections. Our findings expose a concerning paradox: the stronger reasoning capabilities of LLMs increase their vulnerability to reasoning-based poisoning attacks. To mitigate potential misuse, we propose a document isolation-based defense, HODOR, which decouples adversarial logical dependencies.
Jia-Chang Zhang, Min Chen, Xiao-Nan Ren et al.· 0 citations
Continual learning (CL) is a key paradigm that enables intelligent agents to operate autonomously in edge networks over the long term. However, continuous model updates can lead to catastrophic forgetting and representation instability in edge deployment scenarios, which may further induce Decision Boundary Drift (DBD). We propose a DBD-based adversarial attack framework that exploits class-level drift modeling and leverages the deformation of decision boundaries caused by incremental updates. We introduce multiple statistical metrics to quantify boundary drift, based on which class-level adversarial perturbations are constructed and further optimized in the input space to generate effective adversarial examples. Extensive experiments on multiple datasets and continual learning models demonstrate that the proposed method can significantly degrade model robustness, revealing non-negligible security risks in continuously evolving learning systems. Inspired by the security and trustworthiness requirements of edge intelligent agents, we systematically study and quantify DBD and its associated security risks in continual learning. Our findings reveal a practical yet underestimated attack surface and provide a foundation for future research on secure and robust continual learning systems.
Kaixiang Yang, Yue-Bin Xu, Zhi-Hao Li et al.· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.