What it produced
45.5% more records
3,527 records from registries that could not send data, on top of the 7,757 that the other two routes had collected. Counted at the time the paper was published.
Pirmani et al., JMIR Medical Informatics 11(1):e48030, 2023.
Why no single registry was enough
Some of the drugs people with multiple sclerosis take calm the immune system on purpose, and in 2020 nobody knew what that meant once COVID-19 arrived. Doctors had to give advice anyway, and fast.
Registry. A long running collection of records about people with one condition, usually kept by a network of centers or by a country.
The patients who could answer the question were spread over many countries. No single registry had enough COVID-19 cases. So the MS International Federation and the MS Data Alliance opened a worldwide collection, the Global Data Sharing Initiative. It had two ways in: a registry could send its patient records under a signed agreement, or people could fill in a form themselves.
The registries that could not join
Under the two routes that existed, a registry whose rules said the records stay where they are was simply out, and every registry left out meant fewer patients for a question that needed all of them.
The other hard part was agreement. There were eighteen registries under different laws, each with its own custodian and its own IT people, and all of them had to agree on what counted as the same thing before anyone could run anything. Getting that to happen was most of the work.
Moving the analysis instead of the records
Federated model sharing. The analysis is sent to the data, and only its results come back.
I built a third way in. The registry keeps its records, runs the analysis on its own machine and sends back only the results, which are then combined with what came in through the other two routes. The field calls this federated model sharing.
I chose it because the registries with the strictest rules could take part without changing them.
Three ways into one study
Records represented in the study11,284
All three routes. Records from each are represented in the study. Only one route kept the records where they were. Press a route to see what it carried.
Public form. People filled in the questions themselves: 1,383 records from 67 countries. What traveled was their answers.
Record sharing. Registries sent patient records under a signed agreement: 6,374 records from 14 registries. What traveled was the records.
The federated route, the one I built. Registries that were not allowed to send records ran the analysis on their own machine and sent back only the results. It added 45.5% on top of the other two routes.
What we built
A pipeline in three layers, meant for more than this one question.
- First a shared core data set, built from what the researchers needed to answer rather than from what was easy to collect.
- Three ways to collect it: the public form, record sharing and the federated route.
- Quality checks on every stream, then combining them, then the analysis.
Who did what
- Me
- I designed the federated route and the data pipeline, worked with the registries to get it running on their side, and wrote the paper about it.
- The registries
- The four federated registries each ran the analysis on their own systems, and agreed in advance what could leave. The others sent their records under signed agreements.
- The initiative
- The MS International Federation and the MS Data Alliance opened the collection and funded the work.
- Clinical teams
- The medical analyses, and the papers on the combined data.
What changed
The federated route brought in 3,527 records from four registries by the time the paper was published. A fifth registry joined through it later. Together, the three routes built one of the largest data sets of people with multiple sclerosis and COVID-19 at the time.
Analyses of that combined data set informed the worldwide COVID-19 advice for people with multiple sclerosis, as the MS International Federation reports.
What kind of result this is
This is a data access result: how many more patients could take part. The medical findings are in the clinical papers on this data.
The detailFor technical readers
Data and design
- Direct entry. 1,383 records, from people in 67 countries.
- Core data set sharing. 6,374 records, from 14 registries under signed agreements.
- Federated model sharing. 3,527 records, from 4 registries at publication.
- Total. 11,284 records, from 18 registries and the public form.
Core data set. The short list of variables every source agreed to provide, defined the same way everywhere.
The standardization was demand driven: the core data set came from the questions, then each stream was mapped onto it. Every stream went through the same quality step before integration.
Limitations
- Results that come back through the federated route can only be combined in the ways they were designed for. Some analyses that are possible on pooled records are not possible there.
- It ran inside each registry's own systems, so there is no single public repository to point to.
Publication
Pirmani A, De Brouwer E, Geys L, Parciak T, Moreau Y, Peeters LM. The journey of data within a global data sharing initiative: a federated 3-layer data analysis pipeline to scale up multiple sclerosis research. JMIR Medical Informatics 11(1):e48030, 2023. The clinical papers on this data set are on the publications page.