Jul 2026· Data mining and knowledge discovery· Vol 40· 0 citations· 71 references
Computer Science
TL;DR
This work introduces CoLEDS, a method for profiling unlabeled client datasets with minimal computational overhead that yields federatively trained models that are better aligned with individual data distributions and enables appropriate model assignment even for clients that do not participate in federated training.
Abstract
Clustering clients into groups with relatively homogeneous data distributions is a key strategy for improving federated learning under non-independent and identically distributed data. However, most state-of-the-art clustering approaches require clients to possess labeled datasets and perform substantial local computation, limiting their applicability in real-world settings. To address these limitations, we introduce CoLEDS, a method for profiling unlabeled client datasets with minimal computational overhead. CoLEDS trains a model using a contrastive learning objective defined across multiple clients and optimized in a distributed fashion through joint client–server coordination. The resulting model embeds key properties of client datasets into low-dimensional vectors that are shared with the server for clustering. Extensive empirical evaluation shows that these profiles accurately capture latent dataset characteristics. By clustering clients based on these representations, CoLEDS yields federatively trained models that are better aligned with individual data distributions and enables appropriate model assignment even for clients that do not participate in federated training.
SensCluster is proposed, a novel sensitivity-aware CFL framework that constructs compact client representations by selecting parameters that are most responsive to local feature distributions, and consistently outperforms state-of-the-art CFL methods across diverse feature skew scenarios.
Jiaqi Wang, Tobias Schlagenhauf, Setareh Maghsudi· Proceedings of the 32nd ACM...· 0 citations
Local Inference Guided Aggregation for Heterogeneous Training Environments to Yield Enhancement Through Agreement and Regularization (LIGHTYEAR), a federated learning framework that performs update selection in function space using an NTK-based agreement score to characterize predictive behavior and determine a personalized aggregation set for each client.
Mirko Konstantin, S. Zachow, Anirban Mukhopadhyay· 0 citations
Federated learning provides a promising paradigm for collaborative model training among mutually untrusted parties without sharing local data. However, data distributions in real-world federated scenarios are usually heterogeneous, which can significantly degrade global model performance. Existing approaches mainly address the non-independent and identically distributed (Non-IID) problem by optimizing aggregation strategies or sharing auxiliary data. Nevertheless, these methods often exhibit limited effectiveness under severe heterogeneity, insufficient adaptability to highly skewed data distributions, and additional privacy concerns. To address these challenges, this paper proposes AE-FRL, a federated representation learning framework for heterogeneous data. Specifically, AE-FRL employs autoencoders to extract latent representations from local training data and utilizes a representation sharing mechanism to mitigate client drift caused by Non-IID data, thereby improving global model accuracy. To overcome the limited expressiveness of latent representations, a jointly trained supervised autoencoder is introduced, which incorporates downstream classification supervision during sample reconstruction. This design enhances both the discriminative capability and representation quality of the learned latent features. Furthermore, a representation mixup mechanism is proposed to reduce privacy leakage risks during representation sharing and improve robustness against potential inference attacks. Experimantal results on four real-world datasets demonstrate that under various Non-IID settings, AE-FRL consistently outperforms baseline methods including FedAvg, FedProx, SCAFFOLD, FedNova, and FedMix, achieving higher model accuracy. In highly heterogeneous scenarios, AE-FRL achieves over 53.94% accuracy improvement without sacrificing communication efficiency or privacy preservation.
Ling-Tao Tang, Hao-Tian He, Di Wang et al.· 2026 12th International Conf...· 0 citations
Experiments under representative Non-IID settings on benchmark datasets show that PFLS-One achieves improved accuracy and faster convergence compared with representative baseline methods, and the convergence analysis under a non-convex objective provides theoretical support for the proposed method.
Experimental results show that routing-aware collaboration consistently improves personalized performance compared to conventional federated averaging and local training, while maintaining the same communication cost, and shows that client-centric and expert-centric clustering provides an effective and scalable approach for personalized federated instruction fine-tuning of sparse MoE LLMs.
Ankita Sharma, B. Farahani, S. Moosavi et al.· 0 citations
Federated learning enables privacy-preserving collaboration across distributed devices without centralizing local data. However, clients may differ not only in data distributions but also in domain knowledge and annotation capabilities. In this paper, we introduce label granularity skew, a new form of statistical heterogeneity in federated hierarchical classification, in which clients provide taxonomy-consistent labels at different levels of detail within a shared class hierarchy. To model this heterogeneity, we generate client-specific local label hierarchies using a probabilistic relational neighbor classifier and construct a WordNet-guided hierarchy via silhouette score-based coarsening. Our analysis shows that strongly coupled hierarchical models are sensitive to incomplete supervision, while the conditional softmax classifier is more robust. Based on this insight, we propose Branch-wise Decoupled Fine-Tuning (BDFT) and its federated version, FedBDFT, which fine-tune branch-wise classifiers and aggregate them through federated optimization. Experiments on CIFAR-100, TinyImageNet, and ImageNet show that FedBDFT substantially improves robustness under severe label granularity skew, with average gains of 27.9% and 56.4% at skewness levels of 0.6 and 0.9, respectively. Zero-shot results further indicate that FedBDFT better preserves hierarchical representations for unseen fine-grained classes. These findings demonstrate its effectiveness for federated hierarchical classification with heterogeneous label granularities.
Jaeheon Kim, Hokeun Kim, B. Choi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.