IDEAS home Printed from https://ideas.repec.org/a/spr/compst/v36y2021i4d10.1007_s00180-021-01088-1.html
   My bibliography  Save this article

Comparison of the EM, CEM and SEM algorithms in the estimation of finite mixtures of linear mixed models: a simulation study

Author

Listed:
  • Luísa Novais

    (University of Minho)

  • Susana Faria

    (University of Minho)

Abstract

Finite mixture models are a widely known method for modelling data that arise from a heterogeneous population. Within the family of mixtures of regression models, mixtures of linear mixed models have also been applied in different areas since, besides taking into consideration the heterogeneity in the population, they also allow to take into account the correlation between observations from the same individual. One of the main issues in mixture models concerns the estimation of the parameters. Maximum likelihood estimation is one of the most used methods in the estimation of the parameters for mixture models. However, the maximization of the log-likelihood function in mixture models is complex, producing in many cases infinite solutions whereby the maximum likelihood estimator may not exist, at least globally. For this reason, it is common to resort to iterative methods, in particular to the Expectation-Maximization (EM) algorithm. However, the slow convergence and the selection of initial values are two of biggest issues of the EM algorithm, the reason why some modified versions of this algorithm have been developed over the years. In this article we compare the performance of the EM, Classification EM (CEM) and Stochastic EM (SEM) algorithms in the estimation of the parameters for mixtures of linear mixed models. In order to evaluate their performance, we carry out a simulation study and a real data application. The results show that the CEM algorithm is the least computationally demanding algorithm, although the three algorithms provide similar maximum likelihood estimates for the parameters.

Suggested Citation

  • Luísa Novais & Susana Faria, 2021. "Comparison of the EM, CEM and SEM algorithms in the estimation of finite mixtures of linear mixed models: a simulation study," Computational Statistics, Springer, vol. 36(4), pages 2507-2533, December.
  • Handle: RePEc:spr:compst:v:36:y:2021:i:4:d:10.1007_s00180-021-01088-1
    DOI: 10.1007/s00180-021-01088-1
    as

    Download full text from publisher

    File URL: http://link.springer.com/10.1007/s00180-021-01088-1
    File Function: Abstract
    Download Restriction: Access to the full text of the articles in this series is restricted.

    File URL: https://libkey.io/10.1007/s00180-021-01088-1?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Xiao‐Li Meng & David Van Dyk, 1997. "The EM Algorithm—an Old Folk‐song Sung to a Fast New Tune," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 59(3), pages 511-567.
    2. S. Ganesalingam, 1989. "Classification and Mixture Approaches to Clustering Via Maximum Likelihood," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 38(3), pages 455-466, November.
    3. Derek S. Young & David R. Hunter, 2015. "Random effects regression mixtures for analyzing infant habituation," Journal of Applied Statistics, Taylor & Francis Journals, vol. 42(7), pages 1421-1441, July.
    4. Benaglia, Tatiana & Chauveau, Didier & Hunter, David R. & Young, Derek S., 2009. "mixtools: An R Package for Analyzing Mixture Models," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 32(i06).
    5. Yau, Kelvin K. W. & Lee, Andy H. & Ng, Angus S. K., 2003. "Finite mixture regression model with random effects: application to neonatal hospital length of stay," Computational Statistics & Data Analysis, Elsevier, vol. 41(3-4), pages 359-366, January.
    6. Celeux, Gilles & Govaert, Gerard, 1992. "A classification EM algorithm for clustering and two stochastic versions," Computational Statistics & Data Analysis, Elsevier, vol. 14(3), pages 315-332, October.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Addey, Kwame Asiam & Nganje, William, 2024. "Climate policy volatility hinders renewable energy consumption: Evidence from yardstick competition theory," Energy Economics, Elsevier, vol. 130(C).

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Derek S. Young & Xi Chen & Dilrukshi C. Hewage & Ricardo Nilo-Poyanco, 2019. "Finite mixture-of-gamma distributions: estimation, inference, and model-based clustering," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 13(4), pages 1053-1082, December.
    2. Govaert, G. & Nadif, M., 1996. "Comparison of the mixture and the classification maximum likelihood in cluster analysis with binary data," Computational Statistics & Data Analysis, Elsevier, vol. 23(1), pages 65-81, November.
    3. Adrian O’Hagan & Arthur White, 2019. "Improved model-based clustering performance using Bayesian initialization averaging," Computational Statistics, Springer, vol. 34(1), pages 201-231, March.
    4. François Bavaud, 2009. "Aggregation invariance in general clustering approaches," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 3(3), pages 205-225, December.
    5. Wan-Lun Wang, 2019. "Mixture of multivariate t nonlinear mixed models for multiple longitudinal data with heterogeneity and missing values," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(1), pages 196-222, March.
    6. Ana Pinto & Tong Yin & Marion Reichenbach & Raghavendra Bhatta & Pradeep Kumar Malik & Eva Schlecht & Sven König, 2020. "Enteric Methane Emissions of Dairy Cattle Considering Breed Composition, Pasture Management, Housing Conditions and Feeding Characteristics along a Rural-Urban Gradient in a Rising Megacity," Agriculture, MDPI, vol. 10(12), pages 1-18, December.
    7. Rasmus Lentz & Jean Marc Robin & Suphanit Piyapromdee, 2018. "On Worker and Firm Heterogeneity in Wages and Employment Mobility: Evidence from Danish Register Data," 2018 Meeting Papers 469, Society for Economic Dynamics.
    8. Faicel Chamroukhi, 2016. "Piecewise Regression Mixture for Simultaneous Functional Data Clustering and Optimal Segmentation," Journal of Classification, Springer;The Classification Society, vol. 33(3), pages 374-411, October.
    9. Ozonder, Gozde & Miller, Eric J., 2021. "Longitudinal investigation of skeletal activity episode timing decisions – A copula approach," Journal of choice modelling, Elsevier, vol. 40(C).
    10. Mukhopadhyay, Subhadeep & Ghosh, Anil K., 2011. "Bayesian multiscale smoothing in supervised and semi-supervised kernel discriminant analysis," Computational Statistics & Data Analysis, Elsevier, vol. 55(7), pages 2344-2353, July.
    11. Zhou, Lin & Tang, Yayong, 2021. "Linearly preconditioned nonlinear conjugate gradient acceleration of the PX-EM algorithm," Computational Statistics & Data Analysis, Elsevier, vol. 155(C).
    12. Chun Wang & Steven W. Nydick, 2020. "On Longitudinal Item Response Theory Models: A Didactic," Journal of Educational and Behavioral Statistics, , vol. 45(3), pages 339-368, June.
    13. Minjung Kyung & Ju-Hyun Park & Ji Yeh Choi, 2022. "Bayesian Mixture Model of Extended Redundancy Analysis," Psychometrika, Springer;The Psychometric Society, vol. 87(3), pages 946-966, September.
    14. Grün, Bettina & Leisch, Friedrich, 2008. "FlexMix Version 2: Finite Mixtures with Concomitant Variables and Varying and Constant Parameters," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 28(i04).
    15. Xue, Jiacheng & Yao, Weixin, 2022. "Machine Learning Embedded Semiparametric Mixtures of Regressions with Covariate-Varying Mixing Proportions," Econometrics and Statistics, Elsevier, vol. 22(C), pages 159-171.
    16. Meng Li & Sijia Xiang & Weixin Yao, 2016. "Robust estimation of the number of components for mixtures of linear regression models," Computational Statistics, Springer, vol. 31(4), pages 1539-1555, December.
    17. Hornik, Kurt & Grün, Bettina, 2014. "movMF: An R Package for Fitting Mixtures of von Mises-Fisher Distributions," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 58(i10).
    18. Xiong, Yingge & Tobias, Justin L. & Mannering, Fred L., 2014. "The analysis of vehicle crash injury-severity data: A Markov switching approach with road-segment heterogeneity," Transportation Research Part B: Methodological, Elsevier, vol. 67(C), pages 109-128.
    19. Alegre, Joaquín & Mateo, Sara & Pou, Llorenç, 2011. "A latent class approach to tourists’ length of stay," Tourism Management, Elsevier, vol. 32(3), pages 555-563.
    20. M. Vrac & L. Billard & E. Diday & A. Chédin, 2012. "Copula analysis of mixture models," Computational Statistics, Springer, vol. 27(3), pages 427-457, September.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:compst:v:36:y:2021:i:4:d:10.1007_s00180-021-01088-1. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.