IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0059795.html
   My bibliography  Save this article

Accelerating Bayesian Hierarchical Clustering of Time Series Data with a Randomised Algorithm

Author

Listed:
  • Robert Darkins
  • Emma J Cooke
  • Zoubin Ghahramani
  • Paul D W Kirk
  • David L Wild
  • Richard S Savage

Abstract

We live in an era of abundant data. This has necessitated the development of new and innovative statistical algorithms to get the most from experimental data. For example, faster algorithms make practical the analysis of larger genomic data sets, allowing us to extend the utility of cutting-edge statistical methods. We present a randomised algorithm that accelerates the clustering of time series data using the Bayesian Hierarchical Clustering (BHC) statistical method. BHC is a general method for clustering any discretely sampled time series data. In this paper we focus on a particular application to microarray gene expression data. We define and analyse the randomised algorithm, before presenting results on both synthetic and real biological data sets. We show that the randomised algorithm leads to substantial gains in speed with minimal loss in clustering quality. The randomised time series BHC algorithm is available as part of the R package BHC, which is available for download from Bioconductor (version 2.10 and above) via http://bioconductor.org/packages/2.10/bioc/html/BHC.html. We have also made available a set of R scripts which can be used to reproduce the analyses carried out in this paper. These are available from the following URL. https://sites.google.com/site/randomisedbhc/.

Suggested Citation

  • Robert Darkins & Emma J Cooke & Zoubin Ghahramani & Paul D W Kirk & David L Wild & Richard S Savage, 2013. "Accelerating Bayesian Hierarchical Clustering of Time Series Data with a Randomised Algorithm," PLOS ONE, Public Library of Science, vol. 8(4), pages 1-9, April.
  • Handle: RePEc:plo:pone00:0059795
    DOI: 10.1371/journal.pone.0059795
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0059795
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0059795&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0059795?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Brock, Guy & Pihur, Vasyl & Datta, Susmita & Datta, Somnath, 2008. "clValid: An R Package for Cluster Validation," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 25(i04).
    2. Fruhwirth-Schnatter, Sylvia & Kaufmann, Sylvia, 2008. "Model-Based Clustering of Multiple Time Series," Journal of Business & Economic Statistics, American Statistical Association, vol. 26, pages 78-89, January.
    3. L. Bauwens & J. V. K. Rombouts, 2007. "Bayesian Clustering of Many Garch Models," Econometric Reviews, Taylor & Francis Journals, vol. 26(2-4), pages 365-386.
    4. Heard, Nicholas A. & Holmes, Christopher C. & Stephens, David A., 2006. "A Quantitative Study of Gene Regulation Involved in the Immune Response of Anopheline Mosquitoes: An Application of Bayesian Hierarchical Clustering of Curves," Journal of the American Statistical Association, American Statistical Association, vol. 101, pages 18-29, March.
    5. Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Guillaume Marrelec & Arnaud Messé & Pierre Bellec, 2015. "A Bayesian Alternative to Mutual Information for the Hierarchical Clustering of Dependent Random Variables," PLOS ONE, Public Library of Science, vol. 10(9), pages 1-26, September.
    2. Crook Oliver M. & Gatto Laurent & Kirk Paul D. W., 2019. "Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer proteomics," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 18(6), pages 1-20, December.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Jane L. Harvill & Priya Kohli & Nalini Ravishanker, 2017. "Clustering Nonlinear, Nonstationary Time Series Using BSLEX," Methodology and Computing in Applied Probability, Springer, vol. 19(3), pages 935-955, September.
    2. Sylvia Frühwirth-Schnatter, 2011. "Panel data analysis: a survey on model-based clustering of time series," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 5(4), pages 251-280, December.
    3. Ana Alina Tudoran, 2022. "A machine learning approach to identifying decision-making styles for managing customer relationships," Electronic Markets, Springer;IIM University of St. Gallen, vol. 32(1), pages 351-374, March.
    4. Wu, Han-Ming, 2011. "On biological validity indices for soft clustering algorithms for gene expression data," Computational Statistics & Data Analysis, Elsevier, vol. 55(5), pages 1969-1979, May.
    5. Tyler Roick & Dimitris Karlis & Paul D. McNicholas, 2021. "Clustering discrete-valued time series," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 15(1), pages 209-229, March.
    6. De Angelis, Luca & Dias, José G., 2014. "Mining categorical sequences from data using a hybrid clustering method," European Journal of Operational Research, Elsevier, vol. 234(3), pages 720-730.
    7. Volodymyr Melnykov & Xuwen Zhu, 2019. "An extension of the K-means algorithm to clustering skewed data," Computational Statistics, Springer, vol. 34(1), pages 373-394, March.
    8. Edoardo Otranto & Massimo Mucciardi, 2019. "Clustering space-time series: FSTAR as a flexible STAR approach," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 13(1), pages 175-199, March.
    9. Rombouts Jeroen V. K. & Bouaddi Mohammed, 2009. "Mixed Exponential Power Asymmetric Conditional Heteroskedasticity," Studies in Nonlinear Dynamics & Econometrics, De Gruyter, vol. 13(3), pages 1-32, May.
    10. Aßmann, Christian & Boysen-Hogrefe, Jens, 2009. "A bayesian approach to model-based clustering for panel probit models," Economics Working Papers 2009-03, Christian-Albrechts-University of Kiel, Department of Economics.
    11. Aßmann, Christian & Boysen-Hogrefe, Jens, 2011. "A Bayesian approach to model-based clustering for binary panel probit models," Computational Statistics & Data Analysis, Elsevier, vol. 55(1), pages 261-279, January.
    12. E. Otranto & M. Mucciardi, 2017. "Clustering Space-Time Series: A Flexible STAR Approach," Working Paper CRENoS 201707, Centre for North South Economic Research, University of Cagliari and Sassari, Sardinia.
    13. Johann Kraus & Christoph Müssel & Günther Palm & Hans Kestler, 2011. "Multi-objective selection for collecting cluster alternatives," Computational Statistics, Springer, vol. 26(2), pages 341-353, June.
    14. Korsuk Sirinukunwattana & Richard S Savage & Muhammad F Bari & David R J Snead & Nasir M Rajpoot, 2013. "Bayesian Hierarchical Clustering for Studying Cancer Gene Expression Data with Unknown Statistics," PLOS ONE, Public Library of Science, vol. 8(10), pages 1-11, October.
    15. Lu, Emiao & Handl, Julia & Xu, Dong-ling, 2018. "Determining analogies based on the integration of multiple information sources," International Journal of Forecasting, Elsevier, vol. 34(3), pages 507-528.
    16. Juarez, Miguel A. & Steel, Mark F. J., 2006. "Model-based Clustering of non-Gaussian Panel Data," MPRA Paper 880, University Library of Munich, Germany.
    17. Miriam Aparicio, 2021. "Resiliency and Cooperation or Regarding Social and Collective Competencies for University Achievement. An Analysis from a Systemic Perspective," European Journal of Social Sciences Education and Research Articles, Revistia Research and Publishing, vol. 8, ejser_v8_.
    18. Yunpeng Zhao & Qing Pan & Chengan Du, 2019. "Logistic regression augmented community detection for network data with application in identifying autism‐related gene pathways," Biometrics, The International Biometric Society, vol. 75(1), pages 222-234, March.
    19. Cavit Pakel & Neil Shephard & Kevin Sheppard, 2009. "Nuisance parameters, composite likelihoods and a panel of GARCH models," Economics Papers 2009-W12, Economics Group, Nuffield College, University of Oxford.
    20. Patrick Zschech & Kai Heinrich & Raphael Bink & Janis S. Neufeld, 2019. "Prognostic Model Development with Missing Labels," Business & Information Systems Engineering: The International Journal of WIRTSCHAFTSINFORMATIK, Springer;Gesellschaft für Informatik e.V. (GI), vol. 61(3), pages 327-343, June.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0059795. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.