IDEAS home Printed from https://ideas.repec.org/p/cpr/ceprdp/14046.html
   My bibliography  Save this paper

Undiscounted Bandit Games

Author

Listed:
  • Rady, Sven
  • Keller, R Godfrey

Abstract

We analyze undiscounted continuous-time games of strategic experimentation with two-armed bandits. The risky arm generates payoffs according to a Lévy process with an unknown average payoff per unit of time which nature draws from an arbitrary finite set. Observing all actions and realized payoffs, players use Markov strategies with the common posterior belief about the unknown parameter as the state variable. We show that the unique symmetric Markov perfect equilibrium can be computed in a simple closed form involving only the payoff of the safe arm, the expected current payoff of the risky arm, and the expected full-information payoff, given the current belief. In particular, the equilibrium does not depend on the precise specification of the payoff-generating processes.

Suggested Citation

  • Rady, Sven & Keller, R Godfrey, 2019. "Undiscounted Bandit Games," CEPR Discussion Papers 14046, C.E.P.R. Discussion Papers.
  • Handle: RePEc:cpr:ceprdp:14046
    as

    Download full text from publisher

    File URL: https://cepr.org/publications/DP14046
    Download Restriction: CEPR Discussion Papers are free to download for our researchers, subscribers and members. If you fall into one of these categories but have trouble downloading our papers, please contact us at subscribers@cepr.org
    ---><---

    As the access to this document is restricted, you may want to look for a different version below or search for a different version of it.

    Other versions of this item:

    References listed on IDEAS

    as
    1. Asaf Cohen & Eilon Solan, 2013. "Bandit Problems with Lévy Processes," Mathematics of Operations Research, INFORMS, vol. 38(1), pages 92-107, February.
    2. Godfrey Keller & Sven Rady, 1999. "Optimal Experimentation in a Changing Environment," The Review of Economic Studies, Review of Economic Studies Ltd, vol. 66(3), pages 475-507.
    3. Bergemann, Dirk & Valimaki, Juuso, 2002. "Entry and Vertical Differentiation," Journal of Economic Theory, Elsevier, vol. 106(1), pages 91-125, September.
    4. Godfrey Keller & Sven Rady & Martin Cripps, 2005. "Strategic Experimentation with Exponential Bandits," Econometrica, Econometric Society, vol. 73(1), pages 39-68, January.
    5. Patrick Bolton & Christopher Harris, 1999. "Strategic Experimentation," Econometrica, Econometric Society, vol. 67(2), pages 349-374, March.
    6. Dirk Bergemann & Juuso Valimaki, 1997. "Market Diffusion with Two-Sided Learning," RAND Journal of Economics, The RAND Corporation, vol. 28(4), pages 773-795, Winter.
    7. Ke, T. Tony & Villas-Boas, J. Miguel, 2019. "Optimal learning before choice," Journal of Economic Theory, Elsevier, vol. 180(C), pages 383-437.
    8. , & ,, 2010. "Strategic experimentation with Poisson bandits," Theoretical Economics, Econometric Society, vol. 5(2), May.
    9. Martin Peitz & Sven Rady & Piers Trepper, 2017. "Experimentation in Two-Sided Markets," Journal of the European Economic Association, European Economic Association, vol. 15(1), pages 128-172.
    10. Christopher Harris, 1993. "Generalized Solutions of Stochastic Differential Games in One Dimension," Papers 0044, Boston University - Industry Studies Programme.
    11. Pietro Veronesi, 2000. "How Does Information Quality Affect Stock Returns?," Journal of Finance, American Finance Association, vol. 55(2), pages 807-837, April.
    12. Alessandro Bonatti, 2011. "Menu Pricing and Learning," American Economic Journal: Microeconomics, American Economic Association, vol. 3(3), pages 124-163, August.
    13. Dutta, Prajit K., 1991. "What do discounted optima converge to?: A theory of discount rate asymptotics in economic models," Journal of Economic Theory, Elsevier, vol. 55(1), pages 64-94, October.
    14. Moscarini, Giuseppe & Squintani, Francesco, 2010. "Competitive experimentation with private information: The survivor's curse," Journal of Economic Theory, Elsevier, vol. 145(2), pages 639-660, March.
    15. Jovanovic, Boyan, 1979. "Job Matching and the Theory of Turnover," Journal of Political Economy, University of Chicago Press, vol. 87(5), pages 972-990, October.
    16. Keller, Godfrey & Rady, Sven, 2003. "Price Dispersion and Learning in a Dynamic Differentiated-Goods Duopoly," RAND Journal of Economics, The RAND Corporation, vol. 34(1), pages 138-165, Spring.
    17. Dutta, P.K., 1991. "What Do Discounted Optima Converge To? A Theory of Discount Rate Asymptotics in Economic Models," RCER Working Papers 264, University of Rochester - Center for Economic Research (RCER).
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Weng, Xi, 2015. "Dynamic pricing in the presence of individual learning," Journal of Economic Theory, Elsevier, vol. 155(C), pages 262-299.
    2. Martin Peitz & Sven Rady & Piers Trepper, 2017. "Experimentation in Two-Sided Markets," Journal of the European Economic Association, European Economic Association, vol. 15(1), pages 128-172.
    3. Alessandro Bonatti, 2008. "Continuous-Time Screening Contracts," 2008 Meeting Papers 493, Society for Economic Dynamics.
    4. Keller, Godfrey & Rady, Sven, 2015. "Breakdowns," Theoretical Economics, Econometric Society, vol. 10(1), January.
    5. Décamps, Jean-Paul & Mariotti, Thomas & Villeneuve, Stéphane, 2000. "Investment Timing under Incomplete Information," IDEI Working Papers 115, Institut d'Économie Industrielle (IDEI), Toulouse, revised Apr 2004.
    6. Rosenberg, Dinah & Salomon, Antoine & Vieille, Nicolas, 2013. "On games of strategic experimentation," Games and Economic Behavior, Elsevier, vol. 82(C), pages 31-51.
    7. Alessandro Lizzeri & Eran Shmaya & Leeat Yariv, 2024. "Disentangling Exploration from Exploitation," NBER Working Papers 32424, National Bureau of Economic Research, Inc.
    8. Bloch, Francis & Fabrizi, Simona & Lippert, Steffen, 2022. "Hiding and herding in market entry," Journal of Economic Theory, Elsevier, vol. 206(C).
    9. Dinah Rosenberg & Eilon Solan & Nicolas Vieille, 2007. "Social Learning in One-Arm Bandit Problems," Econometrica, Econometric Society, vol. 75(6), pages 1591-1611, November.
    10. Jan Eeckhout & Xi Weng, 2022. "Assortative Learning," Economica, London School of Economics and Political Science, vol. 89(355), pages 647-688, July.
    11. Wagner, Peter A. & Klein, Nicolas, 2022. "Strategic investment and learning with private information," Journal of Economic Theory, Elsevier, vol. 204(C).
    12. Eeckhout, Jan & Weng, Xi, 2015. "Common value experimentation," Journal of Economic Theory, Elsevier, vol. 160(C), pages 317-339.
    13. Godfrey Keller & Sven Rady, 1998. "Market Experimentation in a Dynamic Differentiated-Goods Duopoly," Game Theory and Information 9810001, University Library of Munich, Germany, revised 20 Aug 1999.
    14. Boyarchenko, Svetlana, 2021. "Inefficiency of sponsored research," Journal of Mathematical Economics, Elsevier, vol. 95(C).
    15. Axel Anderson & Luís M. B. Cabral, 2007. "Go for broke or play it safe? Dynamic competition with choice of variance," RAND Journal of Economics, RAND Corporation, vol. 38(3), pages 593-609, September.
    16. Strulovici, Bruno & Szydlowski, Martin, 2015. "On the smoothness of value functions and the existence of optimal strategies in diffusion models," Journal of Economic Theory, Elsevier, vol. 159(PB), pages 1016-1055.
    17. Roland Fryer & Philipp Harms, 2018. "Two-Armed Restless Bandits with Imperfect Information: Stochastic Control and Indexability," Mathematics of Operations Research, INFORMS, vol. 43(2), pages 399-427, May.
    18. Dinah Rosenberg & Eilon Solan & Nicolas Vieille, 2004. "Timing Games with Informational Externalities," Levine's Working Paper Archive 122247000000000704, David K. Levine.
    19. Bonatti, Alessandro & Hörner, Johannes, 2017. "Learning to disagree in a game of experimentation," Journal of Economic Theory, Elsevier, vol. 169(C), pages 234-269.
    20. Keller, Godfrey & Novák, Vladimír & Willems, Tim, 2019. "A note on optimal experimentation under risk aversion," Journal of Economic Theory, Elsevier, vol. 179(C), pages 476-487.

    More about this item

    Keywords

    Strategic experimentation; Two-armed bandit; Strong long-run average criterion; Markov perfect equilibrium; Hjb equation; Viscosity solution;
    All these keywords.

    JEL classification:

    • C73 - Mathematical and Quantitative Methods - - Game Theory and Bargaining Theory - - - Stochastic and Dynamic Games; Evolutionary Games
    • D83 - Microeconomics - - Information, Knowledge, and Uncertainty - - - Search; Learning; Information and Knowledge; Communication; Belief; Unawareness

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:cpr:ceprdp:14046. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: the person in charge (email available below). General contact details of provider: https://www.cepr.org .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.