Federated Learning with Amazon EKS and NVIDIA FLARE: Analysis of Three Data Scenarios
Key point
Global model accuracy improved by 26 percentage points over local models in asymmetric data environments
Details
In environments where data privacy and regulatory compliance are critical, the need for Federated Learning, which trains models without moving raw data, is growing. This content deploys NVIDIA FLARE on Amazon EKS to perform federated learning in a real cluster environment and analyzes its effectiveness and limitations through three data scenarios.
NVIDIA FLARE Architecture Based on Amazon EKS
NVIDIA FLARE is an open-source federated learning framework based on a Python SDK, configured on an Amazon EKS cluster with 1 server and 3 clients. The deployment structure is designed with two tiers—an always-running parent Pod and job Pods dynamically created during tasks—to enhance resource efficiency. The server acts as a controller, assigning tasks to each client, while clients train models on local data and transmit only the weights to the server.
Experimental Results for Three Data Scenarios
The experiments were conducted using the FedAvg algorithm based on MNIST data, verifying changes in global model performance according to data distribution and quality through Cross-Site Evaluation (CSE).
- Uniform Data Distribution (IID): When all three sites had evenly distributed labels from 0 to 9, the global model reached 97% accuracy. This was similar to the performance of local models at each site, indicating limited additional benefits from federated learning.
- Asymmetric Data Distribution: When each site held data excluding specific labels (e.g., 0, 3, 6), local models remained at approximately 70% accuracy. In contrast, the global model recovered to 96% accuracy by compensating for each other's data blind spots, demonstrating the core value of federated learning.
- Inclusion of Low-Quality Data Participants: When one site (site-3) provided only single-label data, resulting in overfitting, that local model's performance collapsed on other sites. However, thanks to sample-count weighting, the global model maintained 98% accuracy, defending against the impact of low-quality data.
Implications
Federated learning demonstrates the ability to build robust global models even in environments where data is siloed. In particular, it can significantly reduce the performance gap compared to local models when data distributions are asymmetric. However, robustness against low-quality data participants does not imply complete defense against intentional attacks; therefore, operators must identify and manage imposter participants through the CSE matrix.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.