Abstract Motivation Federated learning (FL) enables collaborative model training on geographically distributed genomic and clinical datasets while complying with data privacy laws and regulatory constraints. FeatureCloud is an existing platform for FL that provides an accessible web-based interface and a large repository of implemented methods. However, due to its graphical interface, FeatureCloud requires manual interaction of all participants, limiting automation, iteration, and reproducibility. Results We introduce fedflow, a Python-based command-line tool for headless orchestration of FL tasks with FeatureCloud. This tool uses distributed computing resources such as virtual machines or cloud instances to automate such workflows. This allows for scalable federated computing either in local simulations or deployed in a trusted environment. Further, we demonstrate how fedflow can be used to integrate FeatureCloud in reproducible Snakemake workflows. For this, we reanalyse a metagenomic dataset with two federated algorithms and compare the results to the centralized approach with pooled data. Overall, fedflow enables automation of multi-client FL tasks, facilitates embedding of FeatureCloud in standard bioinformatics pipelines and thereby helps increase reproducibility. Availability Fedflow is open-source and available at https://github.com/W-L/fedflow.
Foundation models have demonstrated immense value for scRNA-seq analysis, but their fine-tuning or inference on heterogeneous, privacy-sensitive clinical cohorts is governed by strict data protection policies, which often prohibit centralization. We introduce Clifti-GPT, a privacy-preserving federated framework based on secure multi-party computation (SMPC) that enables collaborative model training and transferable inference, where zero-shot predictions are performed across decentralized clinical repositories by securely aggregating local statistics rather than transferring data embeddings, without sharing patient data, clinical-level statistics, or models. Built upon the scGPT foundation model, Clifti-GPT achieves performance within 4% of centralized scGPT baselines in accuracy, precision, recall, and macro-F1 for cell type classification and reference mapping across six datasets. Furthermore, it demonstrates rapid convergence in terms of communication rounds, reaching 99% of centralized performance on cell type classification in at most two federated rounds on two evaluated datasets, and scales robustly to 30 clients with less than 2% accuracy loss on a large-scale federated cell type classification setting. Our analysis shows that batch effects impact both Clifti-GPT and centralized baseline, while correction leads to similar results across evaluation metrics in heterogeneous settings for both models. Together, these results indicate that Clifti-GPT enables effective fine-tuning and application of single-cell foundation models across distributed clinical datasets in a manner that is GDPR-compatible by design and addresses real-world privacy and institutional data-governance requirements.
Mohammad Bakhtiari, M. Elkjaer, Ali Oğuz Can et al.· BioData Mining· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.