Person: Huang, Xudong
Email Address
AA Acceptance Date
Birth Date
Research Projects
Organizational Units
Job Title
Last Name
First Name
Name
Search Results
Publication Supervised Learning-Based tagSNP Selection for Genome-Wide Disease Classifications
(BioMed Central, 2008) Liu, Qingzhong; Yang, Jack; Chen, Zhongxue; Yang, Mary Qu; Sung, Andrew H; Huang, XudongBackground: Comprehensive evaluation of common genetic variations through association of single nucleotide polymorphisms (SNPs) with complex human diseases on the genome-wide scale is an active area in human genome research. One of the fundamental questions in a SNP-disease association study is to find an optimal subset of SNPs with predicting power for disease status. To find that subset while reducing study burden in terms of time and costs, one can potentially reconcile information redundancy from associations between SNP markers. Results: We have developed a feature selection method named Supervised Recursive Feature Addition (SRFA). This method combines supervised learning and statistical measures for the chosen candidate features/SNPs to reconcile the redundancy information and, in doing so, improve the classification performance in association studies. Additionally, we have proposed a Support Vector based Recursive Feature Addition (SVRFA) scheme in SNP-disease association analysis. Conclusions: We have proposed using SRFA with different statistical learning classifiers and SVRFA for both SNP selection and disease classification and then applying them to two complex disease data sets. In general, our approaches outperform the well-known feature selection method of Support Vector Machine Recursive Feature Elimination and logic regression-based SNP selection for disease classification in genetic association studies. Our study further indicates that both genetic and environmental variables should be taken into account when doing disease predictions and classifications for the most complex human diseases that have gene-environment interactions.
Publication Investigation of Transmembrane Proteins Using a Computational Approach
(BioMed Central, 2008) Yang, Jack Y; Yang, Mary Qu; Dunker, A Keith; Deng, Youping; Huang, XudongBackground: An important subfamily of membrane proteins are the transmembrane α-helical proteins, in which the membrane-spanning regions are made up of α-helices. Given the obvious biological and medical significance of these proteins, it is of tremendous practical importance to identify the location of transmembrane segments. The difficulty of inferring the secondary or tertiary structure of transmembrane proteins using experimental techniques has led to a surge of interest in applying techniques from machine learning and bioinformatics to infer secondary structure from primary structure in these proteins. We are therefore interested in determining which physicochemical properties are most useful for discriminating transmembrane segments from non-transmembrane segments in transmembrane proteins, and for discriminating intrinsically unstructured segments from intrinsically structured segments in transmembrane proteins, and in using the results of these investigations to develop classifiers to identify transmembrane segments in transmembrane proteins. Results: We determined that the most useful properties for discriminating transmembrane segments from non-transmembrane segments and for discriminating intrinsically unstructured segments from intrinsically structured segments in transmembrane proteins were hydropathy, polarity, and flexibility, and used the results of this analysis to construct classifiers to discriminate transmembrane segments from non-transmembrane segments using four classification techniques: two variants of the Self-Organizing Global Ranking algorithm, a decision tree algorithm, and a support vector machine algorithm. All four techniques exhibited good performance, with out-of-sample accuracies of approximately 75%. Conclusions: Several interesting observations emerged from our study: intrinsically unstructured segments and transmembrane segments tend to have opposite properties; transmembrane proteins appear to be much richer in intrinsically unstructured segments than other proteins; and, in approximately 70% of transmembrane proteins that contain intrinsically unstructured segments, the intrinsically unstructured segments are close to transmembrane segments.
Publication Feature Selection and Classification of MAQC-II Breast Cancer and Multiple Myeloma Microarray Gene Expression Data
(Public Library of Science, 2009) Liu, Qingzhong; Sung, Andrew H.; Chen, Zhongxue; Liu, Jianzhong; Deng, Youping; Huang, XudongMicroarray data has a high dimension of variables but available datasets usually have only a small number of samples, thereby making the study of such datasets interesting and challenging. In the task of analyzing microarray data for the purpose of, e.g., predicting gene-disease association, feature selection is very important because it provides a way to handle the high dimensionality by exploiting information redundancy induced by associations among genetic markers. Judicious feature selection in microarray data analysis can result in significant reduction of cost while maintaining or improving the classification or prediction accuracy of learning machines that are employed to sort out the datasets. In this paper, we propose a gene selection method called Recursive Feature Addition (RFA), which combines supervised learning and statistical similarity measures. We compare our method with the following gene selection methods: Support Vector Machine Recursive Feature Elimination (SVMRFE) Leave-One-Out Calculation Sequential Forward Selection (LOOCSFS) Gradient based Leave-one-out Gene Selection (GLGS) To evaluate the performance of these gene selection methods, we employ several popular learning classifiers on the MicroArray Quality Control phase II on predictive modeling (MAQC-II) breast cancer dataset and the MAQC-II multiple myeloma dataset. Experimental results show that gene selection is strictly paired with learning classifier. Overall, our approach outperforms other compared methods. The biological functional analysis based on the MAQC-II breast cancer dataset convinced us to apply our method for phenotype prediction. Additionally, learning classifiers also play important roles in the classification of microarray data and our experimental results indicate that the Nearest Mean Scale Classifier (NMSC) is a good choice due to its prediction reliability and its stability across the three performance measurements: Testing accuracy, MCC values, and AUC errors.
Publication A Hybrid Machine Learning-Based Method for Classifying the Cushing's Syndrome With Comorbid Adrenocortical Lesions
(BioMed Central, 2008) Yang, Jack Y; Yang, Mary Qu; Luo, Zuojie; Li, Jianling; Deng, Youping; Ma, Yan; Huang, XudongBackground: The prognosis for many cancers could be improved dramatically if they could be detected while still at the microscopic disease stage. It follows from a comprehensive statistical analysis that a number of antigens such as hTERT, PCNA and Ki-67 can be considered as cancer markers, while another set of antigens such as P27KIP1 and FHIT are possible markers for normal tissue. Because more than one marker must be considered to obtain a classification of cancer or no cancer, and if cancer, to classify it as malignant, borderline, or benign, we must develop an intelligent decision system that can fullfill such an unmet medical need. Results: We have developed an intelligent decision system using machine learning techniques and markers to characterize tissue as cancerous, non-cancerous or borderline. The system incorporates learning techniques such as variants of support vector machines, neural networks, decision trees, self-organizing feature maps (SOFM) and recursive maximum contrast trees (RMCT). These variants and algorithms we have developed, tend to detect microscopic pathological changes based on features derived from gene expression levels and metabolic profiles. We have also used immunohistochemistry techniques to measure the gene expression profiles from a number of antigens such as cyclin E, P27KIP1, FHIT, Ki-67, PCNA, Bax, Bcl-2, P53, Fas, FasL and hTERT in several particular types of neuroendocrine tumors such as pheochromocytomas, paragangliomas, and the adrenocortical carcinomas (ACC), adenomas (ACA), and hyperplasia (ACH) involved with Cushing's syndrome. We provided statistical evidence that higher expression levels of hTERT, PCNA and Ki-67 etc. are associated with a higher risk that the tumors are malignant or borderline as opposed to benign. We also investigated whether higher expression levels of P27KIP1 and FHIT, etc., are associated with a decreased risk of adrenomedullary tumors. While no significant difference was found between cell-arrest antigens such as P27KIP1 for malignant, borderline, and benign tumors, there was a significant difference between expression levels of such antigens in normal adrenal medulla samples and in adrenomedullary tumors. Conclusions: Our frame work focused on not only different classification schemes and feature selection algorithms, but also ensemble methods such as boosting and bagging in an effort to improve upon the accuracy of the individual classifiers. It is evident that when all sorts of machine learning and statistically learning techniques are combined appropriately into one integrated intelligent medical decision system, the prediction power can be enhanced significantly. This research has many potential applications; it might provide an alternative diagnostic tool and a better understanding of the mechanisms involved in malignant transformation as well as information that is useful for treatment planning and cancer prevention.
Publication Comparison of Feature Selection and Classification for MALDI-MS Data
(BioMed Central, 2009) Liu, Qingzhong; Sung, Andrew H; Qiao, Mengyu; Chen, Zhongxue; Yang, Jack Y; Yang, Mary Qu; Deng, Youping; Huang, XudongIntroduction: In the classification of Mass Spectrometry (MS) proteomics data, peak detection, feature selection, and learning classifiers are critical to classification accuracy. To better understand which methods are more accurate when classifying data, some publicly available peak detection algorithms for Matrix assisted Laser Desorption Ionization Mass Spectrometry (MALDI-MS) data were recently compared; however, the issue of different feature selection methods and different classification models as they relate to classification performance has not been addressed. With the application of intelligent computing, much progress has been made in the development of feature selection methods and learning classifiers for the analysis of high-throughput biological data. The main objective of this paper is to compare the methods of feature selection and different learning classifiers when applied to MALDI-MS data and to provide a subsequent reference for the analysis of MS proteomics data. Results: We compared a well-known method of feature selection, Support Vector Machine Recursive Feature Elimination (SVMRFE), and a recently developed method, Gradient based Leave-one-out Gene Selection (GLGS) that effectively performs microarray data analysis. We also compared several learning classifiers including K-Nearest Neighbor Classifier (KNNC), Naïve Bayes Classifier (NBC), Nearest Mean Scaled Classifier (NMSC), uncorrelated normal based quadratic Bayes Classifier recorded as UDC, Support Vector Machines, and a distance metric learning for Large Margin Nearest Neighbor classifier (LMNN) based on Mahanalobis distance. To compare, we conducted a comprehensive experimental study using three types of MALDI-MS data. Conclusion: Regarding feature selection, SVMRFE outperformed GLGS in classification. As for the learning classifiers, when classification models derived from the best training were compared, SVMs performed the best with respect to the expected testing accuracy. However, the distance metric learning LMNN outperformed SVMs and other classifiers on evaluating the best testing. In such cases, the optimum classification model based on LMNN is worth investigating for future study.
Publication Independent Component Analysis of Alzheimer's DNA Microarray Gene Expression Data
(BioMed Central, 2009) Kong, Wei; Mou, Xiaoyang; Liu, Qingzhong; Chen, Zhongxue; Vanderburg, Charles; Rogers, Jack; Huang, XudongBackground: Gene microarray technology is an effective tool to investigate the simultaneous activity of multiple cellular pathways from hundreds to thousands of genes. However, because data in the colossal amounts generated by DNA microarray technology are usually complex, noisy, high-dimensional, and often hindered by low statistical power, their exploitation is difficult. To overcome these problems, two kinds of unsupervised analysis methods for microarray data: principal component analysis (PCA) and independent component analysis (ICA) have been developed to accomplish the task. PCA projects the data into a new space spanned by the principal components that are mutually orthonormal to each other. The constraint of mutual orthogonality and second-order statistics technique within PCA algorithms, however, may not be applied to the biological systems studied. Extracting and characterizing the most informative features of the biological signals, however, require higher-order statistics. Results: ICA is one of the unsupervised algorithms that can extract higher-order statistical structures from data and has been applied to DNA microarray gene expression data analysis. We performed FastICA method on DNA microarray gene expression data from Alzheimer's disease (AD) hippocampal tissue samples and consequential gene clustering. Experimental results showed that the ICA method can improve the clustering results of AD samples and identify significant genes. More than 50 significant genes with high expression levels in severe AD were extracted, representing immunity-related protein, metal-related protein, membrane protein, lipoprotein, neuropeptide, cytoskeleton protein, cellular binding protein, and ribosomal protein. Within the aforementioned categories, our method also found 37 significant genes with low expression levels. Moreover, it is worth noting that some oncogenes and phosphorylation-related proteins are expressed in low levels. In comparison to the PCA and support vector machine recursive feature elimination (SVM-RFE) methods, which are widely used in microarray data analysis, ICA can identify more AD-related genes. Furthermore, we have validated and identified many genes that are associated with AD pathogenesis. Conclusion: We demonstrated that ICA exploits higher-order statistics to identify gene expression profiles as linear combinations of elementary expression patterns that lead to the construction of potential AD-related pathogenic pathways. Our computing results also validated that the ICA model outperformed PCA and the SVM-RFE method. This report shows that ICA as a microarray data analysis tool can help us to elucidate the molecular taxonomy of AD and other multifactorial and polygenic complex diseases.
Publication High Content Image Analysis for Human H4 Neuroglioma Cells Exposed to CuO Nanoparticles
(BioMed Central, 2007) Li, Fuhai; Zhou, Xiaobo; Zhu, Jinmin; Ma, Jinwen; Huang, Xudong; Wong, Stephen TCBackground High content screening (HCS)-based image analysis is becoming an important and widely used research tool. Capitalizing this technology, ample cellular information can be extracted from the high content cellular images. In this study, an automated, reliable and quantitative cellular image analysis system developed in house has been employed to quantify the toxic responses of human H4 neuroglioma cells exposed to metal oxide nanoparticles. This system has been proved to be an essential tool in our study.Results The cellular images of H4 neuroglioma cells exposed to different concentrations of CuO nanoparticles were sampled using IN Cell Analyzer 1000. A fully automated cellular image analysis system has been developed to perform the image analysis for cell viability. A multiple adaptive thresholding method was used to classify the pixels of the nuclei image into three classes: bright nuclei, dark nuclei, and background. During the development of our image analysis methodology, we have achieved the followings: (1) The Gaussian filtering with proper scale has been applied to the cellular images for generation of a local intensity maximum inside each nucleus; (2) a novel local intensity maxima detection method based on the gradient vector field has been established; and (3) a statistical model based splitting method was proposed to overcome the under segmentation problem. Computational results indicate that 95.9% nuclei can be detected and segmented correctly by the proposed image analysis system.Conclusion The proposed automated image analysis system can effectively segment the images of human H4 neuroglioma cells exposed to CuO nanoparticles. The computational results confirmed our biological finding that human H4 neuroglioma cells had a dose-dependent toxic response to the insult of CuO nanoparticles.
Publication Physiological and Pathological Role of Alpha-Synuclein in Parkinson’s Disease through Iron Mediated Oxidative Stress; The Role of a Putative Iron-Responsive Element
(Molecular Diversity Preservation International (MDPI), 2009) Olivares, David; Huang, Xudong; Branden, Lars; Greig, Nigel H.; Rogers, JackParkinson’s disease (PD) is the second most common progressive neurodegenerative disorder after Alzheimer’s disease (AD) and represents a large health burden to society. Genetic and oxidative risk factors have been proposed as possible causes, but their relative contribution remains unclear. Dysfunction of alpha-synuclein ((\alpha)-syn) has been associated with PD due to its increased presence, together with iron, in Lewy bodies. Brain oxidative damage caused by iron may be partly mediated by (\alpha)-syn oligomerization during PD pathology. Also, (\alpha)-syn gene dosage can cause familial PD and inhibition of its gene expression by blocking translation via a newly identified Iron Responsive Element-like RNA sequence in its 5’-untranslated region may provide a new PD drug target.
Publication Peroxidase Activity of Cyclooxygenase-2 (COX-2) Cross-links []-Amyloid (A[]) and Generates A[]-COX-2 Hetero-oligomers That Are Increased in Alzheimer's Disease
(American Society for Biochemistry & Molecular Biology (ASBMB), 2004-01-14) Nagano, Seiichi; Huang, Xudong; Payton, Sandra M.; Tanzi, Rudolph E.; Bush, Ashley I.; Moir, RobertOxidative stress is associated with the neuropathology of Alzheimer's disease. We have previously shown that human Abeta has the ability to reduce Fe(III) and Cu(II) and produce hydrogen peroxide coupled with these metals, which is correlated with toxicity against primary neuronal cells. Cyclooxygenase (COX)-2 expression is linked to the progression and severity of pathology in AD. COX is a heme-containing enzyme that produces prostaglandins, and the enzyme also possesses peroxidase activity. Here we investigated the possibility of direct interaction between human Abeta and COX-2 being mediated by the peroxidase activity. Human Abeta formed dimers when it was reacted with COX-2 and hydrogen peroxide. Moreover, the peptide formed a cross-linked complex directly with COX-2. Such cross-linking was not observed with rat Abeta, and the sole tyrosine residue specific for human Abeta might therefore be the site of cross-linking. Similar complexes of Abeta and COX-2 were detected in post-mortem brain samples in greater amounts in AD tissue than in age-matched controls. COX-2-mediated cross-linking may inhibit Abeta catabolism and possibly generate toxic intracellular forms of oligomeric Abeta.
Publication Evidence that the β-Amyloid Plaques of Alzheimer's Disease Represent the Redox-silencing and Entombment of Aβ by Zinc
(American Society for Biochemistry & Molecular Biology (ASBMB), 2000-05-08) Cuajungco, Math P.; Goldstein, Lee E.; Nunomura, Akihiko; Smith, Mark A.; Lim, James T.; Atwood, Craig S.; Huang, Xudong; Farrag, Yasser W.; Perry, George; Bush, Ashley I.Abeta binds Zn(2+), Cu(2+), and Fe(3+) in vitro, and these metals are markedly elevated in the neocortex and especially enriched in amyloid plaque deposits of individuals with Alzheimer's disease (AD). Zn(2+) precipitates Abeta in vitro, and Cu(2+) interaction with Abeta promotes its neurotoxicity, correlating with metal reduction and the cell-free generation of H(2)O(2) (Abeta1-42 > Abeta1-40 > ratAbeta1-40). Because Zn(2+) is redox-inert, we studied the possibility that it may play an inhibitory role in H(2)O(2)-mediated Abeta toxicity. In competition to the cytotoxic potentiation caused by coincubation with Cu(2+), Zn(2+) rescued primary cortical and human embryonic kidney 293 cells that were exposed to Abeta1-42, correlating with the effect of Zn(2+) in suppressing Cu(2+)-dependent H(2)O(2) formation from Abeta1-42. Since plaques contain exceptionally high concentrations of Zn(2+), we examined the relationship between oxidation (8-OH guanosine) levels in AD-affected tissue and histological amyloid burden and found a significant negative correlation. These data suggest a protective role for Zn(2+) in AD, where plaques form as the result of a more robust Zn(2+) antioxidant response to the underlying oxidative attack.