IDEAS home Printed from https://ideas.repec.org/h/elg/eechap/19238_22.html
   My bibliography  Save this book chapter

Predicting match outcomes in football by an Ordered Forest estimator

In: A Modern Guide to Sports Economics

Author

Listed:
  • Daniel Goller
  • Michael C. Knaus
  • Michael Lechner
  • Gabriel Okasa

Abstract

Predicting the outcome of football (i.e. soccer) games based on past information is a non-standard predictive task because of the nature of the game outcome, as well as because of the importance of uncertainty (luck and unobservables). The game outcome consists of the scores of the two teams that are usually either collapsed into a goal-difference or further aggregated to reflect whether the game ended as a win for the home or away team, or as a draw. From a statistical perspective, such outcomes have bounded support and, thus, standard linear modelling can be expected to perform poorly. The large amount of uncertainty in the game outcomes due to just luck or due to game- or team-specific unobservables (e.g. hidden injuries of players, etc.) makes it imperative to use prediction methods that fully exploit the potential of the available information, as well as to uncover the uncertainty of a match outcome. The latter is also relevant when interest is not only in single games but also in a league table at the end of the season. Obviously, such league tables should capture the uncertainty for the single games accumulated over a season to be useful guides on what to expect. Recently, machine learning methods have shown their power in all sorts of prediction problems, in particular in situations where the relation of the variables capturing the information used to predict with the target of the prediction, i.e. here the outcome of the game, is non-linear. However, so far there has been only little development in gearing these methods explicitly towards the estimation of the probabilities of ordered outcomes, such as score differences and points, or just wins, draws, and losses. Lechner and Okasa (2019) propose adapting classical random forest estimation, which is known to have excellent predictive performance (e.g. Biau and Scornet (2016), Fernández-Delgado et al. (2014)) to the problem of predicting probabilities of ordered categorical outcomes, such as the win-draw-loss problem of a football game. In this chapter, we use their approach to predict game outcomes of the German Bundesliga 1 (BL1) based on more than ten years' data on game outcomes as well as extensive information about teams, their players, and their environment. These predictions are then used to obtain the final season rankings in a way that reflects and shows the magnitude of the inherent uncertainty of football games.

Suggested Citation

  • Daniel Goller & Michael C. Knaus & Michael Lechner & Gabriel Okasa, 2021. "Predicting match outcomes in football by an Ordered Forest estimator," Chapters, in: Ruud H. Koning & Stefan Kesenne (ed.), A Modern Guide to Sports Economics, chapter 22, pages 335-355, Edward Elgar Publishing.
  • Handle: RePEc:elg:eechap:19238_22
    as

    Download full text from publisher

    File URL: https://www.elgaronline.com/view/edcoll/9781789906523/9781789906523.00026.xml
    Download Restriction: no
    ---><---

    Other versions of this item:

    References listed on IDEAS

    as
    1. Susan Athey & Julie Tibshirani & Stefan Wager, 2016. "Generalized Random Forests," Papers 1610.01271, arXiv.org, revised Apr 2018.
    2. Alex Bryson & Bernd Frick & Rob Simmons, 2013. "The Returns to Scarce Talent," Journal of Sports Economics, , vol. 14(6), pages 606-628, December.
    3. Egon Franck & Stephan Nüesch, 2012. "Talent And/Or Popularity: What Does It Take To Be A Superstar?," Economic Inquiry, Western Economic Association International, vol. 50(1), pages 202-216, January.
    4. Stefan Wager & Susan Athey, 2018. "Estimation and Inference of Heterogeneous Treatment Effects using Random Forests," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 113(523), pages 1228-1242, July.
    5. Gérard Biau & Erwan Scornet, 2016. "A random forest guided tour," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 25(2), pages 197-227, June.
    6. Leitner, Christoph & Zeileis, Achim & Hornik, Kurt, 2010. "Forecasting sports tournaments by ratings of (prob)abilities: A comparison for the EUROÂ 2008," International Journal of Forecasting, Elsevier, vol. 26(3), pages 471-481, July.
    7. Susan Athey, 2018. "The Impact of Machine Learning on Economics," NBER Chapters, in: The Economics of Artificial Intelligence: An Agenda, pages 507-547, National Bureau of Economic Research, Inc.
    8. Rodney J. Paul & Andrew P. Weinbach, 2007. "Does Sportsbook.com Set Pointspreads to Maximize Profits? Tests of the Levitt Model of Sportsbook Behavior," Journal of Prediction Markets, University of Buckingham Press, vol. 1(3), pages 209-218, December.
    9. Groll Andreas & Kneib Thomas & Mayr Andreas & Schauberger Gunther, 2018. "On the dependency of soccer scores – a sparse bivariate Poisson model for the UEFA European football championship 2016," Journal of Quantitative Analysis in Sports, De Gruyter, vol. 14(2), pages 65-79, June.
    10. Gérard Biau & Erwan Scornet, 2016. "Rejoinder on: A random forest guided tour," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 25(2), pages 264-268, June.
    11. Baboota, Rahul & Kaur, Harleen, 2019. "Predictive analysis and modelling football results using machine learning approach for English Premier League," International Journal of Forecasting, Elsevier, vol. 35(2), pages 741-755.
    12. Steven D. Levitt, 2004. "Why are gambling markets organised so differently from financial markets?," Economic Journal, Royal Economic Society, vol. 114(495), pages 223-246, April.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Daniel Goller, 2023. "Analysing a built-in advantage in asymmetric darts contests using causal machine learning," Annals of Operations Research, Springer, vol. 325(1), pages 649-679, June.
    2. Lechner, Michael & Okasa, Gabriel, 2019. "Random Forest Estimation of the Ordered Choice Model," Economics Working Paper Series 1908, University of St. Gallen, School of Economics and Political Science.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Gabriel Okasa, 2022. "Meta-Learners for Estimation of Causal Effects: Finite Sample Cross-Fit Performance," Papers 2201.12692, arXiv.org.
    2. Daniel Boller & Michael Lechner & Gabriel Okasa, 2021. "The Effect of Sport in Online Dating: Evidence from Causal Machine Learning," Papers 2104.04601, arXiv.org.
    3. repec:diw:diwwpp:dp1980 is not listed on IDEAS
    4. Kayo Murakami & Hideki Shimada & Yoshiaki Ushifusa & Takanori Ida, 2022. "Heterogeneous Treatment Effects Of Nudge And Rebate: Causal Machine Learning In A Field Experiment On Electricity Conservation," International Economic Review, Department of Economics, University of Pennsylvania and Osaka University Institute of Social and Economic Research Association, vol. 63(4), pages 1779-1803, November.
    5. Valente, Marica, 2023. "Policy evaluation of waste pricing programs using heterogeneous causal effect estimation," Journal of Environmental Economics and Management, Elsevier, vol. 117(C).
    6. Patrick Krennmair & Timo Schmid, 2022. "Flexible domain prediction using mixed effects random forests," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(5), pages 1865-1894, November.
    7. Michael C Knaus & Michael Lechner & Anthony Strittmatter, 2021. "Machine learning estimation of heterogeneous causal effects: Empirical Monte Carlo evidence," The Econometrics Journal, Royal Economic Society, vol. 24(1), pages 134-161.
    8. Borup, Daniel & Christensen, Bent Jesper & Mühlbach, Nicolaj Søndergaard & Nielsen, Mikkel Slot, 2023. "Targeting predictors in random forest regression," International Journal of Forecasting, Elsevier, vol. 39(2), pages 841-868.
    9. Yiyi Huo & Yingying Fan & Fang Han, 2023. "On the adaptation of causal forests to manifold data," Papers 2311.16486, arXiv.org, revised Dec 2023.
    10. Escribano, Álvaro & Wang, Dandan, 2021. "Mixed random forest, cointegration, and forecasting gasoline prices," International Journal of Forecasting, Elsevier, vol. 37(4), pages 1442-1462.
    11. Yigit Aydede & Jan Ditzen, 2022. "Identifying the regional drivers of influenza-like illness in Nova Scotia with dominance analysis," Papers 2212.06684, arXiv.org.
    12. Blanquero, Rafael & Carrizosa, Emilio & Molero-Río, Cristina & Romero Morales, Dolores, 2020. "Sparsity in optimal randomized classification trees," European Journal of Operational Research, Elsevier, vol. 284(1), pages 255-272.
    13. Max Biggs & Rim Hariss & Georgia Perakis, 2023. "Constrained optimization of objective functions determined from random forests," Production and Operations Management, Production and Operations Management Society, vol. 32(2), pages 397-415, February.
    14. Susan Athey & Julie Tibshirani & Stefan Wager, 2016. "Generalized Random Forests," Papers 1610.01271, arXiv.org, revised Apr 2018.
    15. Blanquero, Rafael & Carrizosa, Emilio & Molero-Río, Cristina & Morales, Dolores Romero, 2022. "On sparse optimal regression trees," European Journal of Operational Research, Elsevier, vol. 299(3), pages 1045-1054.
    16. Pedro Forquesato, 2022. "Who Benefits from Political Connections in Brazilian Municipalities," Papers 2204.09450, arXiv.org.
    17. Emilio Carrizosa & Cristina Molero-Río & Dolores Romero Morales, 2021. "Mathematical optimization in classification and regression trees," TOP: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 29(1), pages 5-33, April.
    18. Nathan Kallus & Xiaojie Mao, 2023. "Stochastic Optimization Forests," Management Science, INFORMS, vol. 69(4), pages 1975-1994, April.
    19. Antoniadis, Anestis & Lambert-Lacroix, Sophie & Poggi, Jean-Michel, 2021. "Random forests for global sensitivity analysis: A selective review," Reliability Engineering and System Safety, Elsevier, vol. 206(C).
    20. Di Fang & Michael R. Thomsen & Rodolfo M. Nayga & Aaron M. Novotny, 2019. "WIC Participation and Relative Quality of Household Food Purchases: Evidence from FoodAPS," Southern Economic Journal, John Wiley & Sons, vol. 86(1), pages 83-105, July.
    21. Zhexiao Lin & Fang Han, 2022. "On regression-adjusted imputation estimators of the average treatment effect," Papers 2212.05424, arXiv.org, revised Jan 2023.

    More about this item

    Keywords

    Economics and Finance;

    JEL classification:

    • Z29 - Other Special Topics - - Sports Economics - - - Other
    • C53 - Mathematical and Quantitative Methods - - Econometric Modeling - - - Forecasting and Prediction Models; Simulation Methods

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:elg:eechap:19238_22. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Darrel McCalla (email available below). General contact details of provider: http://www.e-elgar.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.