Publication: Deduplication on Large Family Datasets
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Thanks to new technology, data collection of family pedigree charts can be instantly collected from cancer clinics. As in common practice, individuals with high risk of cancer are frequently resampled for further checkups such as annual updates. Due to multiple sampling with the lack of a unique identifier, certain databases cannot be very accurate for population based cancer risk research. Deduplication of data is therefore a promising method to produce workable data that allow for data summaries that can not only provide valuable insight of the role of family history in cancer diagnosis but also validate existing models.
This paper works with the Risk Service database, a collection of pedigree charts of over 400,000 families collected from Feb. 2010 to present. While manual checking and deterministic methods work for smaller and cleaner datasets, probabilistic methods are better for larger studies. Unfortunately in such large datasets, the state of art deduplication methods are very computationally expensive. K-means clustering can help divide the problem into smaller portions. We will examine the efficiency of K-means clustering with various distance metrics between family records. We show that even with data with significant number of missing values, we can achieve over 70% accuracy in retaining matching pairs in the same clusters. The distance metric is the most important component for clustering and we discuss various methods that account for family-tree structure and diagnosis of rare diseases.