Privacy-Preserving Cross-Domain Data Mining Modeling Method Based on Federated Learning
Abstract
This paper proposes a privacy-preserving modeling method for cross-domain data mining based on federated learning. Building on the FedAvg framework, the method integrates cross-domain data node modeling, heterogeneous feature semantic mapping, distribution-shift-aware aggregation, and dynamic gradient perturbation into a unified federated mining process. Each data node retains raw samples and field information locally, performs local feature preprocessing and model training, and maps heterogeneous features into a shared latent semantic space to support cross-domain representation alignment without direct data exchange. During global aggregation, node weights are dynamically adjusted according to sample size and distribution offset, reducing the influence of strongly shifted local domains on global parameter updates. Before model updates are uploaded, local gradients are clipped and perturbed with dynamically decaying Gaussian noise to suppress gradient leakage while limiting convergence loss. The model was evaluated on five federated nodes constructed from the UCI Adult and Bank Marketing datasets. Experimental results show that the proposed method achieves an accuracy of 88.73% and an F1 score of 86.41%, improving upon FedAvg by 3.18 and 3.67 percentage points, respectively. Meanwhile, the success rate of gradient inversion attacks is reduced to 12.84%, demonstrating a favorable balance between cross-domain mining performance and privacy protection capability.