IDEAS home Printed from https://ideas.repec.org/p/rut/rutres/201316.html
   My bibliography  Save this paper

Mining Big Data Using Parsimonious Factor and Shrinkage Methods

Author

Listed:
  • Hyun Hak Kim

    (Bank of Korea)

  • Norman Swanson

    (Rutgers University)

Abstract

A number of recent studies in the economics literature have focused on the usefulness of factor models in the context of prediction using "big data". In this paper, our over-arching question is whether such "big data" are useful for modelling low frequency macroeconomic variables such as unemployment, inflation and GDP. In particular, we analyze the predictive benefits associated with the use dimension reducing independent component analysis (ICA) and sparse principal component analysis (SPCA), coupled with a variety of other factor estimation as well as data shrinkage methods, including bagging, boosting, and the elastic net, among others. We do so by carrying out a forecasting "horse-race", involving the estimation of 28 different baseline model types, each constructed using a variety of specification approaches, estimation approaches, and benchmark econometric models; and all used in the prediction of 11 key macroeconomic variables relevant for monetary policy assessment. In many instances, we find that various of our benchmark specifications, including autoregressive (AR) models, AR models with exogenous variables, and (Bayesian) model averaging, do not dominate more complicated nonlinear methods, and that using a combination of factor and other shrinkage methods often yields superior predictions. For example, simple averaging methods are mean square forecast error (MSFE) "best" in only 9 of 33 key cases considered. This is rather surprising new evidence that model averaging methods do not necessarily yield MSFE-best predictions. However, in order to "beat" model averaging methods, including arithmetic mean and Bayesian averaging approaches, we have introduced into our "horse-race" numerous complex new models involve combining complicated factor estimation methods with interesting new forms of shrinkage. For example, SPCA yields MSFE-best prediction models in many cases, particularly when coupled with shrinkage. This result provides strong new evidence of the usefulness of sophisticated factor based forecasting, and therefore, of the use of "big data" in macroeconometric forecasting.

Suggested Citation

  • Hyun Hak Kim & Norman Swanson, 2013. "Mining Big Data Using Parsimonious Factor and Shrinkage Methods," Departmental Working Papers 201316, Rutgers University, Department of Economics.
  • Handle: RePEc:rut:rutres:201316
    as

    Download full text from publisher

    File URL: http://www.sas.rutgers.edu/virtual/snde/wp/2013-16.pdf
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Banerjee, Anindya & Marcellino, Massimiliano & Masten, Igor, 2014. "Forecasting with factor-augmented error correction models," International Journal of Forecasting, Elsevier, vol. 30(3), pages 589-612.
    2. Jushan Bai & Serena Ng, 2002. "Determining the Number of Factors in Approximate Factor Models," Econometrica, Econometric Society, vol. 70(1), pages 191-221, January.
    3. Inoue, Atsushi & Kilian, Lutz, 2008. "How Useful Is Bagging in Forecasting Economic Time Series? A Case Study of U.S. Consumer Price Inflation," Journal of the American Statistical Association, American Statistical Association, vol. 103, pages 511-522, June.
    4. Ravazzolo, F. & van Dijk, D.J.C. & Paap, R. & Franses, Ph.H.B.F., 2006. "Bayesian Model Averaging in the Presence of Structural Breaks," Econometric Institute Research Papers EI 2006-33, Erasmus University Rotterdam, Erasmus School of Economics (ESE), Econometric Institute.
    5. Bai, Jushan & Ng, Serena, 2008. "Forecasting economic time series using targeted predictors," Journal of Econometrics, Elsevier, vol. 146(2), pages 304-317, October.
    6. Stock, James H. & Watson, Mark W., 2006. "Forecasting with Many Predictors," Handbook of Economic Forecasting, in: G. Elliott & C. Granger & A. Timmermann (ed.), Handbook of Economic Forecasting, edition 1, volume 1, chapter 10, pages 515-554, Elsevier.
    7. S. K. Vines, 2000. "Simple principal components," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 49(4), pages 441-451.
    8. Clark, Todd E. & McCracken, Michael W., 2009. "Tests of Equal Predictive Ability With Real-Time Data," Journal of Business & Economic Statistics, American Statistical Association, vol. 27(4), pages 441-454.
    9. Carmen Fernandez & Eduardo Ley & Mark F. J. Steel, 2001. "Model uncertainty in cross-country growth regressions," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 16(5), pages 563-576.
    10. James H. Stock & Mark W. Watson, 2012. "Generalized Shrinkage Methods for Forecasting Using Many Predictors," Journal of Business & Economic Statistics, Taylor & Francis Journals, vol. 30(4), pages 481-493, June.
    11. Connor, Gregory & Korajczyk, Robert A., 1986. "Performance measurement with the arbitrage pricing theory : A new framework for analysis," Journal of Financial Economics, Elsevier, vol. 15(3), pages 373-394, March.
    12. Fernandez, Carmen & Ley, Eduardo & Steel, Mark F. J., 2001. "Benchmark priors for Bayesian model averaging," Journal of Econometrics, Elsevier, vol. 100(2), pages 381-427, February.
    13. Francis X. Diebold & Jose A. Lopez, 1995. "Forecast evaluation and combination," Research Paper 9525, Federal Reserve Bank of New York.
    14. Forni, Mario & Hallin, Marc & Lippi, Marco & Reichlin, Lucrezia, 2005. "The Generalized Dynamic Factor Model: One-Sided Estimation and Forecasting," Journal of the American Statistical Association, American Statistical Association, vol. 100, pages 830-840, September.
    15. Mario Forni & Marc Hallin & Marco Lippi & Lucrezia Reichlin, 2000. "The Generalized Dynamic-Factor Model: Identification And Estimation," The Review of Economics and Statistics, MIT Press, vol. 82(4), pages 540-554, November.
    16. Chow, Gregory C & Lin, An-loh, 1971. "Best Linear Unbiased Interpolation, Distribution, and Extrapolation of Time Series by Related Series," The Review of Economics and Statistics, MIT Press, vol. 53(4), pages 372-375, November.
    17. McCracken, Michael W., 2007. "Asymptotics for out of sample tests of Granger causality," Journal of Econometrics, Elsevier, vol. 140(2), pages 719-752, October.
    18. Connor, Gregory & Korajczyk, Robert A., 1988. "Risk and return in an equilibrium APT : Application of a new test methodology," Journal of Financial Economics, Elsevier, vol. 21(2), pages 255-289, September.
    19. Ming Yuan & Yi Lin, 2007. "On the non‐negative garrotte estimator," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 69(2), pages 143-161, April.
    20. Zou, Hui, 2006. "The Adaptive Lasso and Its Oracle Properties," Journal of the American Statistical Association, American Statistical Association, vol. 101, pages 1418-1429, December.
    21. Todd E. Clark & Michael W. McCracken, 2009. "Improving Forecast Accuracy By Combining Recursive And Rolling Forecasts," International Economic Review, Department of Economics, University of Pennsylvania and Osaka University Institute of Social and Economic Research Association, vol. 50(2), pages 363-395, May.
    22. Clark, Todd E. & McCracken, Michael W., 2001. "Tests of equal forecast accuracy and encompassing for nested models," Journal of Econometrics, Elsevier, vol. 105(1), pages 85-110, November.
    23. McCracken, Michael W., 2004. "Parameter estimation and tests of equal forecast accuracy between non-nested models," International Journal of Forecasting, Elsevier, vol. 20(3), pages 503-514.
    24. James H. Stock & Mark W. Watson, 2005. "Implications of Dynamic Factor Models for VAR Analysis," NBER Working Papers 11467, National Bureau of Economic Research, Inc.
    25. Connor, Gregory & Korajczyk, Robert A, 1993. "A Test for the Number of Factors in an Approximate Factor Model," Journal of Finance, American Finance Association, vol. 48(4), pages 1263-1291, September.
    26. Boivin, Jean & Ng, Serena, 2006. "Are more data always better for factor analysis?," Journal of Econometrics, Elsevier, vol. 132(1), pages 169-194, May.
    27. Diebold, Francis X & Mariano, Roberto S, 2002. "Comparing Predictive Accuracy," Journal of Business & Economic Statistics, American Statistical Association, vol. 20(1), pages 134-144, January.
    28. Michael Artis & Anindya Banerjee & Massimiliano Marcellino, "undated". "Factor forecasts for the UK," Working Papers 203, IGIER (Innocenzo Gasparini Institute for Economic Research), Bocconi University.
    29. Gary Koop & Simon Potter, 2004. "Forecasting in dynamic factor models using Bayesian model averaging," Econometrics Journal, Royal Economic Society, vol. 7(2), pages 550-565, December.
    30. Jean Boivin & Serena Ng, 2005. "Understanding and Comparing Factor-Based Forecasts," International Journal of Central Banking, International Journal of Central Banking, vol. 1(3), December.
    31. Clemen, Robert T., 1989. "Combining forecasts: A review and annotated bibliography," International Journal of Forecasting, Elsevier, vol. 5(4), pages 559-583.
    32. Mc Cracken, Michael W., 2000. "Robust out-of-sample inference," Journal of Econometrics, Elsevier, vol. 99(2), pages 195-223, December.
    33. Hui Zou & Trevor Hastie, 2005. "Addendum: Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(5), pages 768-768, November.
    34. Aiolfi, Marco & Timmermann, Allan, 2006. "Persistence in forecasting performance and conditional combination strategies," Journal of Econometrics, Elsevier, vol. 135(1-2), pages 31-53.
    35. Mark W. Watson & James H. Stock, 2004. "Combination forecasts of output growth in a seven-country data set," Journal of Forecasting, John Wiley & Sons, Ltd., vol. 23(6), pages 405-430.
    36. Stock J.H. & Watson M.W., 2002. "Forecasting Using Principal Components From a Large Number of Predictors," Journal of the American Statistical Association, American Statistical Association, vol. 97, pages 1167-1179, December.
    37. Hui Zou & Trevor Hastie, 2005. "Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(2), pages 301-320, April.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Corradi, Valentina & Swanson, Norman R., 2014. "Testing for structural stability of factor augmented forecasting models," Journal of Econometrics, Elsevier, vol. 182(1), pages 100-118.
    2. Hyun Hak Kim, 2013. "Forecasting Macroeconomic Variables Using Data Dimension Reduction Methods: The Case of Korea," Working Papers 2013-26, Economic Research Institute, Bank of Korea.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Kim, Hyun Hak & Swanson, Norman R., 2018. "Mining big data using parsimonious factor, machine learning, variable selection and shrinkage methods," International Journal of Forecasting, Elsevier, vol. 34(2), pages 339-354.
    2. Kim, Hyun Hak & Swanson, Norman R., 2014. "Forecasting financial and macroeconomic variables using data reduction methods: New empirical evidence," Journal of Econometrics, Elsevier, vol. 178(P2), pages 352-367.
    3. Cheng, Xu & Hansen, Bruce E., 2015. "Forecasting with factor-augmented regression: A frequentist model averaging approach," Journal of Econometrics, Elsevier, vol. 186(2), pages 280-293.
    4. Hyun Hak Kim, 2013. "Forecasting Macroeconomic Variables Using Data Dimension Reduction Methods: The Case of Korea," Working Papers 2013-26, Economic Research Institute, Bank of Korea.
    5. Xu Cheng & Bruce E. Hansen, 2012. "Forecasting with Factor-Augmented Regression: A Frequentist Model Averaging Approach, Second Version," PIER Working Paper Archive 13-061, Penn Institute for Economic Research, Department of Economics, University of Pennsylvania, revised 03 Sep 2013.
    6. Tan, Xueping & Sirichand, Kavita & Vivian, Andrew & Wang, Xinyu, 2022. "Forecasting European carbon returns using dimension reduction techniques: Commodity versus financial fundamentals," International Journal of Forecasting, Elsevier, vol. 38(3), pages 944-969.
    7. Norman R. Swanson & Weiqi Xiong, 2018. "Big data analytics in economics: What have we learned so far, and where should we go from here?," Canadian Journal of Economics/Revue canadienne d'économique, John Wiley & Sons, vol. 51(3), pages 695-746, August.
    8. Karim Barhoumi & Olivier Darné & Laurent Ferrara, 2010. "Are disaggregate data useful for factor analysis in forecasting French GDP?," Journal of Forecasting, John Wiley & Sons, Ltd., vol. 29(1-2), pages 132-144.
    9. Catherine Doz & Peter Fuleky, 2019. "Dynamic Factor Models," Working Papers 2019-4, University of Hawaii Economic Research Organization, University of Hawaii at Manoa.
    10. Karim Barhoumi & Olivier Darné & Laurent Ferrara, 2014. "Dynamic factor models: A review of the literature," OECD Journal: Journal of Business Cycle Measurement and Analysis, OECD Publishing, Centre for International Research on Economic Tendency Surveys, vol. 2013(2), pages 73-107.
    11. Sandra Eickmeier & Christina Ziegler, 2008. "How successful are dynamic factor models at forecasting output and inflation? A meta-analytic approach," Journal of Forecasting, John Wiley & Sons, Ltd., vol. 27(3), pages 237-265.
    12. Lütkepohl, Helmut, 2014. "Structural vector autoregressive analysis in a data rich environment: A survey," SFB 649 Discussion Papers 2014-004, Humboldt University Berlin, Collaborative Research Center 649: Economic Risk.
    13. Smeekes, Stephan & Wijler, Etienne, 2018. "Macroeconomic forecasting using penalized regression methods," International Journal of Forecasting, Elsevier, vol. 34(3), pages 408-430.
    14. Helmut Lütkepohl, 2014. "Structural Vector Autoregressive Analysis in a Data Rich Environment: A Survey," Discussion Papers of DIW Berlin 1351, DIW Berlin, German Institute for Economic Research.
    15. Kihwan Kim & Norman Swanson, 2013. "Diffusion Index Model Specification and Estimation Using Mixed Frequency Datasets," Departmental Working Papers 201315, Rutgers University, Department of Economics.
    16. Nii Ayi Armah & Norman Swanson, 2010. "Seeing Inside the Black Box: Using Diffusion Index Methodology to Construct Factor Proxies in Large Scale Macroeconomic Time Series Environments," Econometric Reviews, Taylor & Francis Journals, vol. 29(5-6), pages 476-510.
    17. Luciani, Matteo, 2014. "Forecasting with approximate dynamic factor models: The role of non-pervasive shocks," International Journal of Forecasting, Elsevier, vol. 30(1), pages 20-29.
    18. Mayr, Johannes, 2010. "Forecasting Macroeconomic Aggregates," Munich Dissertations in Economics 11140, University of Munich, Department of Economics.
    19. Ng, Serena, 2013. "Variable Selection in Predictive Regressions," Handbook of Economic Forecasting, in: G. Elliott & C. Granger & A. Timmermann (ed.), Handbook of Economic Forecasting, edition 1, volume 2, chapter 0, pages 752-789, Elsevier.
    20. Cepni, Oguzhan & Güney, I. Ethem & Swanson, Norman R., 2019. "Nowcasting and forecasting GDP in emerging markets using global financial and macroeconomic diffusion indexes," International Journal of Forecasting, Elsevier, vol. 35(2), pages 555-572.

    More about this item

    Keywords

    prediction; independent component analysis; robust regression; shrinkage; factors;
    All these keywords.

    JEL classification:

    • C32 - Mathematical and Quantitative Methods - - Multiple or Simultaneous Equation Models; Multiple Variables - - - Time-Series Models; Dynamic Quantile Regressions; Dynamic Treatment Effect Models; Diffusion Processes; State Space Models
    • C53 - Mathematical and Quantitative Methods - - Econometric Modeling - - - Forecasting and Prediction Models; Simulation Methods
    • G17 - Financial Economics - - General Financial Markets - - - Financial Forecasting and Simulation

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:rut:rutres:201316. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: the person in charge (email available below). General contact details of provider: https://edirc.repec.org/data/derutus.html .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.