What it produced
Why one model was not enough
Federated learning trains one model across many places without moving the records: each place trains a copy on its own patients and sends back only what the copy learned.
Its weak point is that the places differ. One shared model can end up average everywhere, missing what matters in each country.
Letting each country keep part of the model
I stopped aiming for one model for everyone, and let the model adapt to each country while it still learns from all of them. The field calls this personalized federated learning.
I tested two ways to do it. The first is a new architecture I designed, AdaptiveDualBranchNet: each country keeps part of the model as its own, and only the other part is shared. The second is fine-tuning a shared model on each country's own records. Then I compared both with the usual options: one shared model, each country on its own, and all records pooled in one place.
Try it: can one line fit four countries?
What changed
ROC-AUC. How well a score separates people who will get worse from people who will not. 0.5 is a coin toss, 1 is perfect.
AUC-PR. The same idea, but it rewards finding the few who get worse without flagging many who do not. It suits outcomes where most people stay stable.
Personalization helped. The best ROC-AUC came from my AdaptiveDualBranchNet with FedProx: 0.8398. Fine-tuning came close, at 0.8375. One model trained on all records pooled in one place scored 0.8092. The same federated method without personalization scored 0.7834.
Against that last one, personalization improved ROC-AUC by 7.2% and AUC-PR by 31%, from 0.408 to 0.535. Against pooling, ROC-AUC was 3.8% higher.
In practice, the model ranks patients by risk about as well as published models for similar tasks, as the paper's discussion notes. That alone does not show a clinical benefit. Before it could inform decisions about a single patient, its calibration needs to be fully reported and improved.
What it does not show
The federation was simulated. The records were already in one registry. I split them into country groups, and each group trained on its own, passing only the model between them. It still has to be shown with groups that are really separate.
The gain over pooling comes from the adapting, since the pooled model, with every record in one place, is still one model for everyone.
The detailFor technical readers
Data and design
MSBase. An international registry of people with multiple sclerosis, with records from routine care in many countries.
- Source. The MSBase registry. 26,246 patients, 283,115 clinical episodes, from 146 centers, grouped into 32 country groups.
- Task. Predict disability progression two years ahead.
- Federation. Simulated. Each country group trained as its own client, and only model parameters moved between them.
Methods and evaluation
FedProx. A version of federated averaging that keeps each site's copy from drifting too far from the shared model.
- Baselines. Federated averaging and FedProx with one global model. Each country alone. All records pooled.
- Personalization. A new architecture, AdaptiveDualBranchNet, that shares some parameters and keeps others local. And personalized fine-tuning of the global models. The best score, 0.8398, is AdaptiveDualBranchNet with FedProx. With federated averaging the same architecture reached 0.8384. Fine-tuning reached 0.8375 with FedProx and 0.8370 with federated averaging.
- Metrics. ROC-AUC and AUC-PR, over ten repetitions.
Code and publication
Pirmani A, De Brouwer E, Arany A, et al. Personalized federated learning for predicting disability progression in multiple sclerosis using real-world routine clinical data. npj Digital Medicine 8(478), 2025.
The study code is public: FL-MS-RWD on GitHub. Python, PyTorch and Flower, run on a high performance computing cluster, because the data agreement required all computing to happen in one approved place.