Hadley, DexterPan, JamesEl-Sayed, OsamaAljabban, JihadAljabban, ImadAzad, Tej D.Hadied, Mohamad O.Raza, ShuaibRayikanti, Benjamin AbhishekChen, BinPaik, HyojungAran, DvirSpatz, JordanHimmelstein, DanielPanahiazar, MaryamBhattacharya, SanchitaSirota, MarinaMusen, Mark A.Butte, Atul J.2017-12-052017Hadley, D., J. Pan, O. El-Sayed, J. Aljabban, I. Aljabban, T. D. Azad, M. O. Hadied, et al. 2017. “Precision annotation of digital samples in NCBI’s gene expression omnibus.” Scientific Data 4 (1): 170125. doi:10.1038/sdata.2017.125. http://dx.doi.org/10.1038/sdata.2017.125.http://nrs.harvard.edu/urn-3:HUL.InstRepos:34491872The Gene Expression Omnibus (GEO) contains more than two million digital samples from functional genomics experiments amassed over almost two decades. However, individual sample meta-data remains poorly described by unstructured free text attributes preventing its largescale reanalysis. We introduce the Search Tag Analyze Resource for GEO as a web application (http://STARGEO.org) to curate better annotations of sample phenotypes uniformly across different studies, and to use these sample annotations to define robust genomic signatures of disease pathology by meta-analysis. In this paper, we target a small group of biomedical graduate students to show rapid crowd-curation of precise sample annotations across all phenotypes, and we demonstrate the biological validity of these crowd-curated annotations for breast cancer. STARGEO.org makes GEO data findable, accessible, interoperable and reusable (i.e., FAIR) to ultimately facilitate knowledge discovery. Our work demonstrates the utility of crowd-curation and interpretation of open ‘big data’ under FAIR principles as a first step towards realizing an ideal paradigm of precision medicine.en-USData miningData acquisitionPrecision annotation of digital samples in NCBI’s gene expression omnibusJournal Article2017-12-0510.1038/sdata.2017.125