IDEAS home Printed from https://ideas.repec.org/a/bla/scjsta/v48y2021i3p729-760.html
   My bibliography  Save this article

Clustering with statistical error control

Author

Listed:
  • Michael Vogt
  • Matthias Schmid

Abstract

This article presents a clustering approach that allows for rigorous statistical error control similar to a statistical test. We develop estimators for both the unknown number of clusters and the clusters themselves. The estimators depend on a tuning parameter α which is similar to the significance level of a statistical hypothesis test. By choosing α, one can control the probability of overestimating the true number of clusters, while the probability of underestimation is asymptotically negligible. In addition, the probability that the estimated clusters differ from the true ones is controlled. In the theoretical part of the article, formal versions of these statements on statistical error control are derived in a baseline model with convex clusters. A simulation study and two applications to temperature and gene expression microarray data complement the theoretical analysis.

Suggested Citation

  • Michael Vogt & Matthias Schmid, 2021. "Clustering with statistical error control," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 48(3), pages 729-760, September.
  • Handle: RePEc:bla:scjsta:v:48:y:2021:i:3:p:729-760
    DOI: 10.1111/sjos.12450
    as

    Download full text from publisher

    File URL: https://doi.org/10.1111/sjos.12450
    Download Restriction: no

    File URL: https://libkey.io/10.1111/sjos.12450?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Inder Tecuapetla-Gómez & Axel Munk, 2017. "Autocovariance Estimation in Regression with a Discontinuous Signal and m-Dependent Errors: A Difference-Based Approach," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 44(2), pages 346-368, June.
    2. Jiahua Chen & Pengfei Li & Yuejiao Fu, 2012. "Inference on the Order of a Normal Mixture," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 107(499), pages 1096-1105, September.
    3. Robert Tibshirani & Guenther Walther & Trevor Hastie, 2001. "Estimating the number of clusters in a data set via the gap statistic," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 63(2), pages 411-423.
    4. Li, Pengfei & Chen, Jiahua, 2010. "Testing the Order of a Finite Mixture," Journal of the American Statistical Association, American Statistical Association, vol. 105(491), pages 1084-1092.
    5. Ranjan Maitra & Volodymyr Melnykov & Soumendra N. Lahiri, 2012. "Bootstrapping for Significance of Compact Clusters in Multidimensional Datasets," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 107(497), pages 378-392, March.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Holzmann, Hajo & Schwaiger, Florian, 2016. "Testing for the number of states in hidden Markov models," Computational Statistics & Data Analysis, Elsevier, vol. 100(C), pages 318-330.
    2. Kasahara Hiroyuki & Shimotsu Katsumi, 2012. "Testing the Number of Components in Finite Mixture Models," Global COE Hi-Stat Discussion Paper Series gd12-259, Institute of Economic Research, Hitotsubashi University.
    3. Hiroyuki Kasahara & Katsumi Shimotsu, 2017. "Testing the Order of Multivariate Normal Mixture Models," CIRJE F-Series CIRJE-F-1044, CIRJE, Faculty of Economics, University of Tokyo.
    4. Wichitchan, Supawadee & Yao, Weixin & Yang, Guangren, 2019. "Hypothesis testing for finite mixture models," Computational Statistics & Data Analysis, Elsevier, vol. 132(C), pages 180-189.
    5. Yanyuan Ma & Shaoli Wang & Lin Xu & Weixin Yao, 2021. "Semiparametric mixture regression with unspecified error distributions," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 30(2), pages 429-444, June.
    6. Maciejowska, Katarzyna, 2013. "Assessing the number of components in a normal mixture: an alternative approach," MPRA Paper 50303, University Library of Munich, Germany.
    7. Bagkavos, Dimitrios & Patil, Prakash N., 2023. "Goodness-of-fit testing for normal mixture densities," Computational Statistics & Data Analysis, Elsevier, vol. 188(C).
    8. Erika S. Helgeson & David M. Vock & Eric Bair, 2021. "Nonparametric cluster significance testing with reference to a unimodal null distribution," Biometrics, The International Biometric Society, vol. 77(4), pages 1215-1226, December.
    9. Xu Gao & Weining Shen & Jing Ning & Ziding Feng & Jianhua Hu, 2022. "Addressing patient heterogeneity in disease predictive model development," Biometrics, The International Biometric Society, vol. 78(3), pages 1045-1055, September.
    10. Thiemo Fetzer & Samuel Marden, 2017. "Take What You Can: Property Rights, Contestability and Conflict," Economic Journal, Royal Economic Society, vol. 0(601), pages 757-783, May.
    11. Khanh Duong, 2024. "Is meritocracy just? New evidence from Boolean analysis and Machine learning," Journal of Computational Social Science, Springer, vol. 7(2), pages 1795-1821, October.
    12. Daniel Agness & Travis Baseler & Sylvain Chassang & Pascaline Dupas & Erik Snowberg, 2022. "Valuing the Time of the Self-Employed," Working Papers 2022-2, Princeton University. Economics Department..
    13. Batool, Fatima & Hennig, Christian, 2021. "Clustering with the Average Silhouette Width," Computational Statistics & Data Analysis, Elsevier, vol. 158(C).
    14. Nicoleta Serban & Huijing Jiang, 2012. "Multilevel Functional Clustering Analysis," Biometrics, The International Biometric Society, vol. 68(3), pages 805-814, September.
    15. Orietta Nicolis & Jean Paul Maidana & Fabian Contreras & Danilo Leal, 2024. "Analyzing the Impact of COVID-19 on Economic Sustainability: A Clustering Approach," Sustainability, MDPI, vol. 16(4), pages 1-30, February.
    16. Li, Pai-Ling & Chiou, Jeng-Min, 2011. "Identifying cluster number for subspace projected functional data clustering," Computational Statistics & Data Analysis, Elsevier, vol. 55(6), pages 2090-2103, June.
    17. Yaeji Lim & Hee-Seok Oh & Ying Kuen Cheung, 2019. "Multiscale Clustering for Functional Data," Journal of Classification, Springer;The Classification Society, vol. 36(2), pages 368-391, July.
    18. Forzani, Liliana & Gieco, Antonella & Tolmasky, Carlos, 2017. "Likelihood ratio test for partial sphericity in high and ultra-high dimensions," Journal of Multivariate Analysis, Elsevier, vol. 159(C), pages 18-38.
    19. Ye, Mao & Lu, Zhao-Hua & Li, Yimei & Song, Xinyuan, 2019. "Finite mixture of varying coefficient model: Estimation and component selection," Journal of Multivariate Analysis, Elsevier, vol. 171(C), pages 452-474.
    20. Yujia Li & Xiangrui Zeng & Chien‐Wei Lin & George C. Tseng, 2022. "Simultaneous estimation of cluster number and feature sparsity in high‐dimensional cluster analysis," Biometrics, The International Biometric Society, vol. 78(2), pages 574-585, June.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bla:scjsta:v:48:y:2021:i:3:p:729-760. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Wiley Content Delivery (email available below). General contact details of provider: http://www.blackwellpublishing.com/journal.asp?ref=0303-6898 .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.