Permitted by design? Article 10(5) AIA, sensitive data, and the legal illusion of bias correction
Permitted by design? Article 10(5) AIA, sensitive data, and the legal illusion of bias correction
- Research Article
- 10.3233/shti260176
- May 21, 2026
- Studies in health technology and informatics
Multi-site medical studies require methods that address confounding bias while respecting data privacy regulations. We propose a federated learning method integrating propensity score matching (PSM) to achieve both objectives simultaneously. Using data from five intensive care unit (ICU) databases (N=160,752), we evaluated within-hospital and cross-hospital PSM strategies within federated XGBoost. Cross-hospital PSM achieved 76.3% mean bias reduction (standardized mean difference (SMD): 0.070), while within-hospital PSM achieved 74.3% reduction (SMD: 0.074), compared to baseline SMD of 0.316. Both strategies maintained predictive performance (AUROC: 0.73-0.75). Our method enables rigorous causal inference in distributed healthcare networks without centralizing sensitive patient data.
- Research Article
16
- 10.1177/1403494819890784
- Dec 11, 2019
- Scandinavian Journal of Public Health
Aims: Selective participation may hamper the validity of population-based cohort studies. The resulting bias can be alleviated by linking auxiliary register data to both the participants and the non-participants of the study, estimating propensity scores for participation and correcting for participation based on these. However, registry holders may not be allowed to disclose sensitive data on (invited) non-participants. Our aim is to provide guidance on how adequate bias correction can be achieved by using auxiliary register data but without disclosing information that could be linked to the subset of non-participants. Methods: We show how existing methods can be used to estimate generalisation weights under various data disclosure scenarios where invited non-participants are indistinguishable from uninvited ones. We also demonstrate how the methods can be implemented using Nordic register data. Results: Inverse-probability-of-sampling weights estimated within a random sample of the target population in which the non-respondents are disclosed are equivalent in expectation to analogous weights in a scenario where the non-participants and uninvited individuals from the population are indistinguishable. To minimise the risk of disclosure when the entire population is invited to participate, investigators should instead consider inverse-odds-of-sampling weights, a method that has previously been suggested for transporting study results to external populations. Conclusions: Generalisation weights can be estimated from auxiliary register data without disclosing information on invited non-participants.