Publication:

Evaluating the Statistical Fidelity and Utility of Aggregated Patient Count as a Surrogate for Line-Level EHR in Kidney-Transplant Analyses

Loading...
Thumbnail Image

Date

2026-05-15

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Chen, Yanshi. 2026. Evaluating the Statistical Fidelity and Utility of Aggregated Patient Count as a Surrogate for Line-Level EHR in Kidney-Transplant Analyses. Masters Thesis, Harvard Medical School.

Abstract

Electronic Health Records (EHR) contain valuable information for clinical and population-level studies, but the sharing of raw data is strictly constrained by privacy regulations. Cumulus addresses this challenge by generating privacy-preserving aggregated patient-count data, termed “CUBE”, which is a power set matrix of variable combinations with k-anonymity constraints. However, the statistical fidelity and utility of CUBE data compared to traditional line-level EHR remains unquantified. We conducted evaluations comparing CUBE and line-level EHR using data from a cohort of patients with kidney transplant at the Boston Children’s Hospital. Our approach includes: (1) implement Bayesian inference to model hidden information in the CUBE, (2) evaluate the statistical fidelity between the original data and the CUBE, and (3) evaluate the utility of CUBE in kidney-transplant-related analyses relative to using the original data. The count inference pipeline outperformed baseline imputation strategies in recovering suppressed counts, achieving mean absolute error (MAE) of 0.12 in a high-dimensional, low-cardinality CUBE (10 variables, 57.1% of suppressed cells deterministically recovered) and MAE of 1.53 in a lower-dimensional, high-cardinality CUBE (3 variables). Our results from the fidelity evaluation demonstrated that the CUBE preserved marginal and joint distributions with consistently low Jensen-Shannon divergence and retains pairwise associations with low mean absolute difference of Cramér’s V against the original data. The results from the utility evaluation demonstrated that the CUBE reproduced conclusions consistent with the original data in descriptive statistics, odds ratios, and classification performance. These findings show that CUBE data retains sufficient statistical fidelity and utility for clinical research applications.

Description

Other Available Sources

Research Data

Keywords

Bioinformatics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories