Effects of Non-IID Distributions in Lung Cancer Data on Survival Prediction with Federated Ensemble Learning.
Summary
This study investigates the impact of non-independent and identically distributed (Non-IID) data on the performance of Federated Learning (FL) models for survival prediction using lung cancer data. Researchers compared Random Forest (RF) and AdaBoost algorithms across various scenarios of data distribution shifts among clients. The findings indicate that both algorithms experienced significant performance degradation under highly non-IID conditions. Notably, Random Forest consistently outperformed AdaBoost in these challenging scenarios. The developed FL algorithm holds potential for model personalization and fine-tuning, with broader applicability to other clinical datasets.
Analysis
This research is crucial for the application of federated learning in oncology, where data privacy and patient population diversity are significant challenges. Demonstrating that Random Forest is more robust than AdaBoost against non-IID distributions provides valuable guidance for algorithm selection in real-world clinical settings. This could accelerate the development of more reliable and personalized survival prediction models for lung cancer patients, by leveraging distributed data without compromising privacy.