2021· International Journal of Data Engineering and Intelligent Computing· 0 citations
TL;DR
This paper presents a comprehensive architectural framework for scalable data engineering tailored for FL in heterogeneous and resource-constrained environments and proposes a modular architecture that addresses data heterogeneity, scalability, and compliance.
Abstract
The growing adoption of Federated Learning (FL) is reshaping the way machine learning models are trained across distributed, privacy-sensitive datasets. However, the scalable and efficient orchestration of data engineering pipelines in decentralized cloud environments remains a significant challenge. This paper presents a comprehensive architectural framework for scalable data engineering tailored for FL in heterogeneous and resource-constrained environments. By integrating modern distributed computing paradigms, such as Kubernetes-based orchestration, edge-aware data preprocessing, and secure federated communication, we propose a modular architecture that addresses data heterogeneity, scalability, and compliance. A case study in a healthcare IoT scenario validates the performance and flexibility of the proposed system. Our work serves as a blueprint for deploying robust FL systems in real-world decentralized cloud ecosystems.
Abstract Motivation Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. Results We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. Availability Fedflow is open-source and available at https://github.com/W-L/fedflow.
The findings reveal that combining FL with LLMs can significantly improve security, trust, and operational efficiency in multi-cloud AI collaborations.
Naveen Kumar, M. Krishnan· International Journal of Dat...· 0 citations
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
This paper proposes Federated MLOps, a framework that combines CI/CD principles with federated model training to enable secure, automated, and efficient deployment of distributed ML models.
Ibrahim Yusuf, Amina Bello· International Journal of Mac...· 0 citations
This paper investigates the implementation of FL in distributed cloud systems, highlighting its role in preserving data privacy and improving scalability, and analyzes various FL algorithms, such as Federated Averaging (FedAvg), assessing their effectiveness in edge computing contexts.
Kenji Sato· International Journal of Art...· 0 citations
A review of federated learning through a structured taxonomy that covers its core architectural paradigms, major learning types, model training approaches, and aggregation mechanisms, and analyzes the principal challenges confronting FL, including privacy and security risks, statistical and system heterogeneity, communication constraints, and global model divergence.
Mahdiyeh Velaei, Hosna Ghahramani, Ali Ghaffari et al.· Cluster Computing· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.