Beamforming-Aided Wireless Multimodal Federated Learning: Dynamic Pruning and Joint Optimization
Abstract
Wireless multimodal federated learning (MFL) is promising for privacy-preserving edge intelligence, but its practical deployment is challenged by modality heterogeneity, client resource heterogeneity, and the high communication and computation cost of multimodal models. This paper proposes a beamforming-aided wireless MFL framework with dynamic neural network pruning, termed mFedDNP. In mFedDNP, each scheduled client trains a pruned unimodal submodel matched to its communication and computing capability, while the server uses modal compensation to handle missing modalities and maintains updates in a storage-efficient manner. We characterize the effect of pruning on communication and computation cost, and formulate a joint design over client scheduling, neural network pruning, and resource block (RB) allocation under perround latency constraints. We further develop a staleness-aware solution that combines adaptive pruning with matching-based RB allocation. Experimental results on CREMA-D and UCI-HAR show that mFedDNP consistently outperforms benchmark methods, achieving up to 4.89 percentage points higher Macro-F1 and 23.52% shorter convergence time.