IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v53y2009i6p2264-2274.html
   My bibliography  Save this article

Robust PCA for skewed data and its outlier map

Author

Listed:
  • Hubert, Mia
  • Rousseeuw, Peter
  • Verdonck, Tim

Abstract

The outlier sensitivity of classical principal component analysis (PCA) has spurred the development of robust techniques. Existing robust PCA methods like ROBPCA work best if the non-outlying data have an approximately symmetric distribution. When the original variables are skewed, too many points tend to be flagged as outlying. A robust PCA method is developed which is also suitable for skewed data. To flag the outliers a new outlier map is defined. Its performance is illustrated on real data from economics, engineering, and finance, and confirmed by a simulation study.

Suggested Citation

  • Hubert, Mia & Rousseeuw, Peter & Verdonck, Tim, 2009. "Robust PCA for skewed data and its outlier map," Computational Statistics & Data Analysis, Elsevier, vol. 53(6), pages 2264-2274, April.
  • Handle: RePEc:eee:csdana:v:53:y:2009:i:6:p:2264-2274
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167-9473(08)00287-9
    Download Restriction: Full text for ScienceDirect subscribers only.
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Croux, Christophe & Ruiz-Gazen, Anne, 2005. "High breakdown estimators for principal components: the projection-pursuit approach revisited," Journal of Multivariate Analysis, Elsevier, vol. 95(1), pages 206-226, July.
    2. Hubert, Mia & Engelen, Sanne, 2007. "Fast cross-validation of high-breakdown resampling methods for PCA," Computational Statistics & Data Analysis, Elsevier, vol. 51(10), pages 5013-5024, June.
    3. Serneels, Sven & Verdonck, Tim, 2008. "Principal component analysis for data containing outliers and missing elements," Computational Statistics & Data Analysis, Elsevier, vol. 52(3), pages 1712-1727, January.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Boudt, Kris & Croux, Christophe, 2010. "Robust M-estimation of multivariate GARCH models," Computational Statistics & Data Analysis, Elsevier, vol. 54(11), pages 2459-2469, November.
    2. Osipenko, Maria, 2021. "Directional assessment of traffic flow extremes," Transportation Research Part B: Methodological, Elsevier, vol. 150(C), pages 353-369.
    3. Verpoorten Marijke, 2012. "The Intensity of the Rwandan Genocide: Measures from the Gacaca Records," Peace Economics, Peace Science, and Public Policy, De Gruyter, vol. 18(1), pages 1-26, April.
    4. Debruyne, Michiel & Hubert, Mia & Van Horebeek, Johan, 2010. "Detecting influential observations in Kernel PCA," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3007-3019, December.
    5. Huang, Xiaolin & Shi, Lei & Pelckmans, Kristiaan & Suykens, Johan A.K., 2014. "Asymmetric ν-tube support vector regression," Computational Statistics & Data Analysis, Elsevier, vol. 77(C), pages 371-382.
    6. Iaci, Ross & Sriram, T.N., 2013. "Robust multivariate association and dimension reduction using density divergences," Journal of Multivariate Analysis, Elsevier, vol. 117(C), pages 281-295.
    7. Marianna Succurro, 2017. "Financial Bankruptcy across European Countries," International Journal of Economics and Finance, Canadian Center of Science and Education, vol. 9(7), pages 132-146, July.
    8. Boente, Graciela & Pires, Ana M. & Rodrigues, Isabel M., 2010. "Detecting influential observations in principal components and common principal components," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 2967-2975, December.
    9. repec:lic:licosd:25610 is not listed on IDEAS
    10. Václav Plevka & Pieter Segaert & Chris M. J. Tampère & Mia Hubert, 2016. "Analysis of travel activity determinants using robust statistics," Transportation, Springer, vol. 43(6), pages 979-996, November.
    11. Szafranek, Karol, 2021. "Evidence on time-varying inflation synchronization," Economic Modelling, Elsevier, vol. 94(C), pages 1-13.
    12. Stephane Heritier & Maria-Pia Victoria-Feser, 2018. "Discussion of “The power of monitoring: how to make the most of a contaminated multivariate sample” by Andrea Cerioli, Marco Riani, Anthony C. Atkinson and Aldo Corbellini," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 27(4), pages 595-602, December.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Sven Serneels, 2019. "Projection pursuit based generalized betas accounting for higher order co-moment effects in financial market analysis," Papers 1908.00141, arXiv.org.
    2. Debruyne, Michiel & Hubert, Mia & Van Horebeek, Johan, 2010. "Detecting influential observations in Kernel PCA," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3007-3019, December.
    3. Serneels, Sven & Verdonck, Tim, 2009. "Principal component regression for data containing outliers and missing elements," Computational Statistics & Data Analysis, Elsevier, vol. 53(11), pages 3855-3863, September.
    4. Cevallos-Valdiviezo, Holger & Van Aelst, Stefan, 2019. "Fast computation of robust subspace estimators," Computational Statistics & Data Analysis, Elsevier, vol. 134(C), pages 171-185.
    5. B. Barış Alkan, 2016. "Robust Principal Component Analysis Based on Modified Minimum Covariance Determinant in the Presence of Outliers," Alphanumeric Journal, Bahadir Fatih Yildirim, vol. 4(2), pages 85-94, September.
    6. Graciela Boente & Frank Critchley & Liliana Orellana, 2007. "Influence functions of two families of robust estimators under proportional scatter matrices," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 15(3), pages 295-327, February.
    7. García-Escudero, L.A. & Gordaliza, A. & Mayo-Iscar, A. & San Martín, R., 2010. "Robust clusterwise linear regression through trimming," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3057-3069, December.
    8. Boente, Graciela & Molina, Julieta & Sued, Mariela, 2010. "On the asymptotic behavior of general projection-pursuit estimators under the common principal components model," Statistics & Probability Letters, Elsevier, vol. 80(3-4), pages 228-235, February.
    9. Nengsih Titin Agustin & Bertrand Frédéric & Maumy-Bertrand Myriam & Meyer Nicolas, 2019. "Determining the number of components in PLS regression on incomplete data set," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 18(6), pages 1-28, December.
    10. Heinrich Fritz & Peter Filzmoser & Christophe Croux, 2012. "A comparison of algorithms for the multivariate L 1 -median," Computational Statistics, Springer, vol. 27(3), pages 393-410, September.
    11. Boente, Graciela & Pires, Ana M. & Rodrigues, Isabel M., 2010. "Detecting influential observations in principal components and common principal components," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 2967-2975, December.
    12. Kondylis, Athanassios & Hadi, Ali S., 2006. "Derived components regression using the BACON algorithm," Computational Statistics & Data Analysis, Elsevier, vol. 51(2), pages 556-569, November.
    13. Lanius, Vivian & Gather, Ursula, 2010. "Robust online signal extraction from multivariate time series," Computational Statistics & Data Analysis, Elsevier, vol. 54(4), pages 966-975, April.
    14. Graciela Boente & Matías Salibian-Barrera, 2015. "S -Estimators for Functional Principal Component Analysis," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 110(511), pages 1100-1111, September.
    15. Frahm, Gabriel & Nordhausen, Klaus & Oja, Hannu, 2020. "M-estimation with incomplete and dependent multivariate data," Journal of Multivariate Analysis, Elsevier, vol. 176(C).
    16. Choulakian, V. & Allard, J. & Almhana, J., 2006. "Robust centroid method," Computational Statistics & Data Analysis, Elsevier, vol. 51(2), pages 737-746, November.
    17. Todorov, Valentin & Filzmoser, Peter, 2009. "An Object-Oriented Framework for Robust Multivariate Analysis," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 32(i03).
    18. Brooks, J.P. & Dulá, J.H. & Boone, E.L., 2013. "A pure L1-norm principal component analysis," Computational Statistics & Data Analysis, Elsevier, vol. 61(C), pages 83-98.
    19. Kalogridis, Ioannis & Van Aelst, Stefan, 2019. "Robust functional regression based on principal components," Journal of Multivariate Analysis, Elsevier, vol. 173(C), pages 393-415.
    20. Hyndman, Rob J. & Shahid Ullah, Md., 2007. "Robust forecasting of mortality and fertility rates: A functional data approach," Computational Statistics & Data Analysis, Elsevier, vol. 51(10), pages 4942-4956, June.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:53:y:2009:i:6:p:2264-2274. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.