Work / Disability progression in multiple sclerosis

Disability progression

Two years ahead, will this person get worse?

Multiple sclerosis does not progress the same way in everyone, and a doctor who can see who is likely to get worse in the next two years can act earlier. The records to learn this exist, across many countries. The catch is that patients differ from country to country, and one shared model can miss exactly those differences.

Question
Can routine clinical records, from countries that differ, predict who will get worse?
Constraint
Patients, care and record keeping differ from one country to the next.
What I built
AdaptiveDualBranchNet, a new architecture. I conceived the study, ran the experiments and led the writing, on a paper with 73 authors.
What changed
AUC-PR went from 0.41 to 0.53 against the same method without adapting. Best ROC-AUC 0.8398, against 0.8092 with all records pooled. In simulated federation.

The decisionFederated learning usually trains one model for everyone. I let each of the 32 countries keep part of the model as its own.

2025, npj Digital Medicine. MSBase registry, 26,246 patients from 146 centers. Funded by the Flanders AI Research Program.

On this page
  1. Why one model was not enough
  2. A model that adapts
  3. What changed
  4. What it does not show
  5. Data and design
  6. Methods and evaluation
  7. Code and publication

What it produced

ROC-AUC, SAME PATIENTS, THREE WAYSPersonalized to each country0.8398All records pooled in one place0.8092One shared model, not personalized0.78340.780.800.820.840.86ROC-AUC. The axis starts at 0.76.
ROC-AUC, where 0.5 is a coin toss and 1 is perfect. The axis starts at 0.76 so the gaps are visible. All three use FedProx where federated. Standard deviations over ten repetitions: 0.0019, 0.0012 and 0.0019. From table 1 of the paper.

Why one model was not enough

Federated learning trains one model across many places without moving the records: each place trains a copy on its own patients and sends back only what the copy learned.

Its weak point is that the places differ. One shared model can end up average everywhere, missing what matters in each country.

Letting each country keep part of the model

I stopped aiming for one model for everyone, and let the model adapt to each country while it still learns from all of them. The field calls this personalized federated learning.

I tested two ways to do it. The first is a new architecture I designed, AdaptiveDualBranchNet: each country keeps part of the model as its own, and only the other part is shared. The second is fine-tuning a shared model on each country's own records. Then I compared both with the usual options: one shared model, each country on its own, and all records pooled in one place.

Try it: can one line fit four countries?

Drag the ends of the line, or use the arrow keys. A sketch with made-up points. In the study, personalizing to each country reached a ROC-AUC of 0.8398, against 0.7834 without it.

What changed

ROC-AUC. How well a score separates people who will get worse from people who will not. 0.5 is a coin toss, 1 is perfect.

AUC-PR. The same idea, but it rewards finding the few who get worse without flagging many who do not. It suits outcomes where most people stay stable.

Personalization helped. The best ROC-AUC came from my AdaptiveDualBranchNet with FedProx: 0.8398. Fine-tuning came close, at 0.8375. One model trained on all records pooled in one place scored 0.8092. The same federated method without personalization scored 0.7834.

Against that last one, personalization improved ROC-AUC by 7.2% and AUC-PR by 31%, from 0.408 to 0.535. Against pooling, ROC-AUC was 3.8% higher.

In practice, the model ranks patients by risk about as well as published models for similar tasks, as the paper's discussion notes. That alone does not show a clinical benefit. Before it could inform decisions about a single patient, its calibration needs to be fully reported and improved.

What it does not show

The federation was simulated. The records were already in one registry. I split them into country groups, and each group trained on its own, passing only the model between them. It still has to be shown with groups that are really separate.

The gain over pooling comes from the adapting, since the pooled model, with every record in one place, is still one model for everyone.

The detailFor technical readers

Data and design

MSBase. An international registry of people with multiple sclerosis, with records from routine care in many countries.

  • Source. The MSBase registry. 26,246 patients, 283,115 clinical episodes, from 146 centers, grouped into 32 country groups.
  • Task. Predict disability progression two years ahead.
  • Federation. Simulated. Each country group trained as its own client, and only model parameters moved between them.
One dot per person: 26,246 people, sorted into 32 country groups. The groups are drawn equal here; the real ones differ in size.

Methods and evaluation

FedProx. A version of federated averaging that keeps each site's copy from drifting too far from the shared model.

  • Baselines. Federated averaging and FedProx with one global model. Each country alone. All records pooled.
  • Personalization. A new architecture, AdaptiveDualBranchNet, that shares some parameters and keeps others local. And personalized fine-tuning of the global models. The best score, 0.8398, is AdaptiveDualBranchNet with FedProx. With federated averaging the same architecture reached 0.8384. Fine-tuning reached 0.8375 with FedProx and 0.8370 with federated averaging.
  • Metrics. ROC-AUC and AUC-PR, over ten repetitions.

Code and publication

Pirmani A, De Brouwer E, Arany A, et al. Personalized federated learning for predicting disability progression in multiple sclerosis using real-world routine clinical data. npj Digital Medicine 8(478), 2025.

The study code is public: FL-MS-RWD on GitHub. Python, PyTorch and Flower, run on a high performance computing cluster, because the data agreement required all computing to happen in one approved place.

Let's talk.

If your question depends on health data that can't simply be put in one place, I probably want to hear about it.

Email me

Opens your email app with the question in it. Nothing is sent from here.

TurkishMerhabaAzeri, my mother tongue, as written in Urmia: xoş gəldin, welcomeخوش گلدینAzerbaijani, as written in AzerbaijanSalamPersian, the language of school: dorudدرودEnglishHelloDutch, as said in FlandersDagyou
A greeting in each language I speak, where it lives.