Skip to content

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

Jul 2026 · arXiv.org · Vol abs/2607.26933 · 0 citations · 47 references
Computer Science

TL;DR

FedDAB, a two-phase method that combines local contrastive regularization with alignment checking, to defend against backdoor attacks, is proposed and theoretically proves FedDAB's robustness with a convergence rate of $\mathcal{O}(1/T)$.

Abstract

Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthiness of backdoor attacks. To tackle these issues, we propose FedDAB, a two-phase method that combines local contrastive regularization with alignment checking, to defend against backdoor attacks. In the first phase, FedDAB introduces a novel model-contrastive term into the local objective to enhance direction and magnitude consistency among benign updates. In the second phase, FedDAB employs an alignment checking strategy to evaluate each local update in terms of overall-direction alignment and parameter-level alignment with historical information, excluding updates that exhibit abnormal alignment patterns from global aggregation. We theoretically prove FedDAB's robustness with a convergence rate of $\mathcal{O}(1/T)$. Extensive experiments show that FedDAB outperforms existing defense methods against backdoor attacks.

View source

Similar papers

Conference 2026

FedRGD: Risk-Guided Dynamic Defense against Federated Backdoors

FedRGD is a federated risk-guided dynamic defense framework that enables efficient fine-grained protection against backdoor attacks in non-IID environments, and combines feature inconsistency detection with lightweight masking and robust aggregation to achieve both accuracy and efficiency.

Rui-Ying Wang · 0 citations
2026

PREFed: An Effective and Stealthy Static-Anchor Backdoor Attack via Trigger Pre-Optimization in Federated Learning

Existing Federated Learning (FL) backdoor attacks commonly employ round-wise proximity strategies, dynamically adapting malicious updates to mimic benign ones in order to evade detection. However, such adaptive mechanisms often introduce instability, increase computational overhead, and create temporal patterns that make attacks more detectable. This work presents a theoretical analysis of how attack configurations affect the disparity between benign and malicious model updates. We derive a two-sided bound on the parameter divergence between benign and backdoored local models, characterizing both an upper bound that governs detectability under defense, and a matching lower bound that exposes an irreducible label-flip signal no trigger optimization can eliminate. Guided by these insights, we propose PREFed, a static-anchor backdoor attack framework that leverages the clean data distribution to optimize trigger patterns under standard training configurations. PREFed eliminates the need for round-wise adaptation by pre-optimizing triggers before training, effectively reducing local training overhead and enhancing attack stability and stealth. Comprehensive evaluations on image classification benchmarks demonstrate that PREFed consistently outperforms three state-of-the-art attacks across six advanced defense mechanisms; cross-domain experiments on SST-2 further confirm the generality of the framework. It achieves over 80% backdoor accuracy within five communication rounds while reducing main task accuracy by less than 2%, compared to more than 15% degradation in prior methods. These results validate PREFed as an efficient and stealthy backdoor attack paradigm for practical federated learning environments.

Xi Chen, Rui Zeng, Chun-Yi Zhou et al. · 0 citations
Preprint Aug 2026

BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated Learning

BackDFL is presented, a unified benchmark for systematically evaluating DFL under realistic and adaptive backdoor attacks, and demonstrates that both state-of-the-art Byzantine-robust DFL methods and adapted FL backdoor defenses fail under modest malicious participation rates, especially in heterogeneous settings.

M. Bouchiha, Gregory Blanc, Yu-Fei Han · 0 citations
Open access 2026

Fed-CBE: Client-Side Backdoor Elimination in Federated Learning via Persistent Parameter Disruption

Fed-CBE is proposed, a novel client-side defense algorithm that eliminates backdoors through three synergistic mechanisms: periodic alternating layer resetting disrupts deep parameters to dismantle cross-round backdoor accumulation, and indiscriminate forgetting employs entropy maximization on non-ground-truth classes to decouple backdoor associations without prior trigger knowledge.

Chun-Hai Li, Yun-Hui Shen, Ming Xie et al. · 0 citations
Open access 2026

SiftFL: Scheduling-Based Robust Backdoor Detection in Federated Learning

Backdoor attacks pose a serious threat to federated learning, particularly when client data are non-IID and the attacker ratio is high. FilterFL is a recent server-side defense that employs two Conditional Generative Adversarial Networks (CGANs) to generate synthetic samples and identify malicious client models without requiring clean server data. However, executing both CGAN stages in every communication round makes the defense robust but computationally expensive. In this paper, we propose SiftFL, a scheduling-based robust backdoor detection method that sifts out malicious client models at a fraction of the original cost. SiftFL decouples the cost of the CGAN stages from the number of communication rounds by executing them periodically rather than every round and complements this schedule with a trust history score that stabilizes client filtering across rounds. This design preserves and, in several settings, improves the robustness of CGAN-based detection while sharply lowering its server-side cost. Experiments using MNIST, CIFAR-10, and GTSRB benchmark dataset show that SiftFL reduces server defense computation by up to 99% while keeping the drop in main accuracy within about 4% in the most challenging non-IID cases compared to the original baseline. At the same time, the attack success rate is reduced by roughly 97–99%, and robustness accuracy improves significantly to 85%, in settings where the original FilterFL becomes unstable. The results indicate that scheduling and trust history make SiftFL a more practical and reliable backdoor detection method under non-IID data distribution and high attacker presence.

Ekhlas Hashem, Muhamad Felemban, Sajjad Mahmood et al. · 0 citations
Open access Aug 2026

A Byzantine-Resilient Federated Learning Framework with Cryptographic Gradient Attestation Against Coordinated Model Poisoning Attacks

FedSentinel is presented, a novel Byzantine-resilient federated learning framework that combines cryptographic gradient attestation with adaptive trust-weighted aggregation to protect against coordinated model-poisoning attacks, which are among the most serious challenges.

Abdullah Abdulkarim Alnajim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.