Efficient genotype compression and analysis of large genetic variation datasets
Layer, Ryan M.
Quinlan, Aaron R.
MetadataShow full item record
CitationLayer, Ryan M., Neil Kindlon, Konrad J. Karczewski, and Aaron R. Quinlan. 2015. “Efficient genotype compression and analysis of large genetic variation datasets.” Nature methods 13 (1): 63-65. doi:10.1038/nmeth.3654. http://dx.doi.org/10.1038/nmeth.3654.
AbstractGenotype Query Tools (GQT) is a new indexing strategy that expedites analyses of genome variation datasets in VCF format based on sample genotypes, phenotypes and relationships. GQT’s compressed genotype index minimizes decompression for analysis, and performance relative to existing methods improves with cohort size. We show substantial (up to 443 fold) performance gains over existing methods and demonstrate GQT’s utility for exploring massive datasets involving thousands to millions of genomes.
Citable link to this pagehttp://nrs.harvard.edu/urn-3:HUL.InstRepos:27320241
- HMS Scholarly Articles