IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0040996.html
   My bibliography  Save this article

Constructing Endophenotypes of Complex Diseases Using Non-Negative Matrix Factorization and Adjusted Rand Index

Author

Listed:
  • Hui-Min Wang
  • Ching-Lin Hsiao
  • Ai-Ru Hsieh
  • Ying-Chao Lin
  • Cathy S J Fann

Abstract

Complex diseases are typically caused by combinations of molecular disturbances that vary widely among different patients. Endophenotypes, a combination of genetic factors associated with a disease, offer a simplified approach to dissect complex trait by reducing genetic heterogeneity. Because molecular dissimilarities often exist between patients with indistinguishable disease symptoms, these unique molecular features may reflect pathogenic heterogeneity. To detect molecular dissimilarities among patients and reduce the complexity of high-dimension data, we have explored an endophenotype-identification analytical procedure that combines non-negative matrix factorization (NMF) and adjusted rand index (ARI), a measure of the similarity of two clusterings of a data set. To evaluate this procedure, we compared it with a commonly used method, principal component analysis with k-means clustering (PCA-K). A simulation study with gene expression dataset and genotype information was conducted to examine the performance of our procedure and PCA-K. The results showed that NMF mostly outperformed PCA-K. Additionally, we applied our endophenotype-identification analytical procedure to a publicly available dataset containing data derived from patients with late-onset Alzheimer’s disease (LOAD). NMF distilled information associated with 1,116 transcripts into three metagenes and three molecular subtypes (MS) for patients in the LOAD dataset: MS1 (), MS2 (), and MS3 (). ARI was then used to determine the most representative transcripts for each metagene; 123, 89, and 71 metagene-specific transcripts were identified for MS1, MS2, and MS3, respectively. These metagene-specific transcripts were identified as the endophenotypes. Our results showed that 14, 38, 0, and 28 candidate susceptibility genes listed in AlzGene database were found by all patients, MS1, MS2, and MS3, respectively. Moreover, we found that MS2 might be a normal-like subtype. Our proposed procedure provides an alternative approach to investigate the pathogenic mechanism of disease and better understand the relationship between phenotype and genotype.

Suggested Citation

  • Hui-Min Wang & Ching-Lin Hsiao & Ai-Ru Hsieh & Ying-Chao Lin & Cathy S J Fann, 2012. "Constructing Endophenotypes of Complex Diseases Using Non-Negative Matrix Factorization and Adjusted Rand Index," PLOS ONE, Public Library of Science, vol. 7(7), pages 1-12, July.
  • Handle: RePEc:plo:pone00:0040996
    DOI: 10.1371/journal.pone.0040996
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0040996
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0040996&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0040996?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Vivian G. Cheung & Richard S. Spielman & Kathryn G. Ewens & Teresa M. Weber & Michael Morley & Joshua T. Burdick, 2005. "Mapping determinants of human gene expression by regional and genome-wide association," Nature, Nature, vol. 437(7063), pages 1365-1369, October.
    2. Karthik Devarajan, 2008. "Nonnegative Matrix Factorization: An Analytical and Interpretive Tool in Computational Biology," PLOS Computational Biology, Public Library of Science, vol. 4(7), pages 1-12, July.
    3. Daniel D. Lee & H. Sebastian Seung, 1999. "Learning the parts of objects by non-negative matrix factorization," Nature, Nature, vol. 401(6755), pages 788-791, October.
    4. Allison, David B. & Gadbury, Gary L. & Heo, Moonseong & Fernandez, Jose R. & Lee, Cheol-Koo & Prolla, Tomas A. & Weindruch, Richard, 2002. "A mixture model approach for the analysis of microarray gene expression data," Computational Statistics & Data Analysis, Elsevier, vol. 39(1), pages 1-20, March.
    5. Yanqing Chen & Jun Zhu & Pek Yee Lum & Xia Yang & Shirly Pinto & Douglas J. MacNeil & Chunsheng Zhang & John Lamb & Stephen Edwards & Solveig K. Sieberts & Amy Leonardson & Lawrence W. Castellini & Su, 2008. "Variations in DNA elucidate molecular networks that cause disease," Nature, Nature, vol. 452(7186), pages 429-435, March.
    6. Yujin Hoshida & Jean-Philippe Brunet & Pablo Tamayo & Todd R Golub & Jill P Mesirov, 2007. "Subclass Mapping: Identifying Common Subtypes in Independent Disease Data Sets," PLOS ONE, Public Library of Science, vol. 2(11), pages 1-8, November.
    7. Michael Morley & Cliona M. Molony & Teresa M. Weber & James L. Devlin & Kathryn G. Ewens & Richard S. Spielman & Vivian G. Cheung, 2004. "Genetic analysis of genome-wide variation in human gene expression," Nature, Nature, vol. 430(7001), pages 743-747, August.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Cordelia Ziraldo & Yoram Vodovotz & Rami A Namas & Khalid Almahmoud & Victor Tapias & Qi Mi & Derek Barclay & Bahiyyah S Jefferson & Guoqiang Chen & Timothy R Billiar & Ruben Zamora, 2013. "Central Role for MCP-1/CCL2 in Injury-Induced Inflammation Revealed by In Vitro, In Silico, and Clinical Studies," PLOS ONE, Public Library of Science, vol. 8(12), pages 1-18, December.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Jin Hyun Ju & Sushila A Shenoy & Ronald G Crystal & Jason G Mezey, 2017. "An independent component analysis confounding factor correction framework for identifying broad impact expression quantitative trait loci," PLOS Computational Biology, Public Library of Science, vol. 13(5), pages 1-26, May.
    2. Paul Fogel & Yann Gaston-Mathé & Douglas Hawkins & Fajwel Fogel & George Luta & S. Stanley Young, 2016. "Applications of a Novel Clustering Approach Using Non-Negative Matrix Factorization to Environmental Research in Public Health," IJERPH, MDPI, vol. 13(5), pages 1-14, May.
    3. Yixin Fang & Yang Feng & Ming Yuan, 2014. "Regularized principal components of heritability," Computational Statistics, Springer, vol. 29(3), pages 455-465, June.
    4. Flavia Esposito, 2021. "A Review on Initialization Methods for Nonnegative Matrix Factorization: Towards Omics Data Experiments," Mathematics, MDPI, vol. 9(9), pages 1-17, April.
    5. Barbara E Stranger & Stephen B Montgomery & Antigone S Dimas & Leopold Parts & Oliver Stegle & Catherine E Ingle & Magda Sekowska & George Davey Smith & David Evans & Maria Gutierrez-Arcelus & Alkes P, 2012. "Patterns of Cis Regulatory Variation in Diverse Human Populations," PLOS Genetics, Public Library of Science, vol. 8(4), pages 1-13, April.
    6. Ryan Abo & Gregory D Jenkins & Liewei Wang & Brooke L Fridley, 2012. "Identifying the Genetic Variation of Gene Expression Using Gene Sets: Application of Novel Gene Set eQTL Approach to PharmGKB and KEGG," PLOS ONE, Public Library of Science, vol. 7(8), pages 1-11, August.
    7. Jingu Kim & Yunlong He & Haesun Park, 2014. "Algorithms for nonnegative matrix and tensor factorizations: a unified view based on block coordinate descent framework," Journal of Global Optimization, Springer, vol. 58(2), pages 285-319, February.
    8. José M. Maisog & Andrew T. DeMarco & Karthik Devarajan & Stanley Young & Paul Fogel & George Luta, 2021. "Assessing Methods for Evaluating the Number of Components in Non-Negative Matrix Factorization," Mathematics, MDPI, vol. 9(22), pages 1-13, November.
    9. Ning Jiang & Minghui Wang & Tianye Jia & Lin Wang & Lindsey Leach & Christine Hackett & David Marshall & Zewei Luo, 2011. "A Robust Statistical Method for Association-Based eQTL Analysis," PLOS ONE, Public Library of Science, vol. 6(8), pages 1-11, August.
    10. Paul C Boutros & Ivy D Moffat & Allan B Okey & Raimo Pohjanvirta, 2011. "mRNA Levels in Control Rat Liver Display Strain-Specific, Hereditary, and AHR-Dependent Components," PLOS ONE, Public Library of Science, vol. 6(7), pages 1-15, July.
    11. Josine L Min & Jennifer M Taylor & J Brent Richards & Tim Watts & Fredrik H Pettersson & John Broxholme & Kourosh R Ahmadi & Gabriela L Surdulescu & Ernesto Lowy & Christian Gieger & Chris Newton-Cheh, 2011. "The Use of Genome-Wide eQTL Associations in Lymphoblastoid Cell Lines to Identify Novel Genetic Pathways Involved in Complex Traits," PLOS ONE, Public Library of Science, vol. 6(7), pages 1-14, July.
    12. Wei Zhang & Jun Zhu & Eric E Schadt & Jun S Liu, 2010. "A Bayesian Partition Method for Detecting Pleiotropic and Epistatic eQTL Modules," PLOS Computational Biology, Public Library of Science, vol. 6(1), pages 1-10, January.
    13. Rafael Teixeira & Mário Antunes & Diogo Gomes & Rui L. Aguiar, 2024. "Comparison of Semantic Similarity Models on Constrained Scenarios," Information Systems Frontiers, Springer, vol. 26(4), pages 1307-1330, August.
    14. Del Corso, Gianna M. & Romani, Francesco, 2019. "Adaptive nonnegative matrix factorization and measure comparisons for recommender systems," Applied Mathematics and Computation, Elsevier, vol. 354(C), pages 164-179.
    15. P Fogel & C Geissler & P Cotte & G Luta, 2022. "Applying separative non-negative matrix factorization to extra-financial data," Working Papers hal-03689774, HAL.
    16. Xiao-Bai Li & Jialun Qin, 2017. "Anonymizing and Sharing Medical Text Records," Information Systems Research, INFORMS, vol. 28(2), pages 332-352, June.
    17. Parrish, Rudolph S. & Spencer III, Horace J. & Xu, Ping, 2009. "Distribution modeling and simulation of gene expression data," Computational Statistics & Data Analysis, Elsevier, vol. 53(5), pages 1650-1660, March.
    18. Julia Schröder & Vitalia Schüller & Andrea May & Christian Gerges & Mario Anders & Jessica Becker & Timo Hess & Nicole Kreuser & René Thieme & Kerstin U Ludwig & Tania Noder & Marino Venerito & Lothar, 2019. "Identification of loci of functional relevance to Barrett’s esophagus and esophageal adenocarcinoma: Cross-referencing of expression quantitative trait loci data from disease-relevant tissues with gen," PLOS ONE, Public Library of Science, vol. 14(12), pages 1-12, December.
    19. Naiyang Guan & Lei Wei & Zhigang Luo & Dacheng Tao, 2013. "Limited-Memory Fast Gradient Descent Method for Graph Regularized Nonnegative Matrix Factorization," PLOS ONE, Public Library of Science, vol. 8(10), pages 1-10, October.
    20. Spelta, A. & Pecora, N. & Rovira Kaltwasser, P., 2019. "Identifying Systemically Important Banks: A temporal approach for macroprudential policies," Journal of Policy Modeling, Elsevier, vol. 41(1), pages 197-218.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0040996. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.