Experimenting on Markov Decision Processes with Local Treatments

My bibliography Save this paper

Experimenting on Markov Decision Processes with Local Treatments

Author

Listed:

Shuze Chen
David Simchi-Levi
Chonghuan Wang

Registered:

Abstract

Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial practice. However, as service systems become increasingly dynamical and personalized, much focus is shifting toward maximizing long-term cumulative outcomes, such as customer lifetime value, through lifetime exposure to interventions. To bridge this gap, we investigate the randomized experiments within dynamical systems modeled as Markov Decision Processes (MDPs). Our goal is to assess the impact of treatment and control policies on long-term cumulative rewards from relatively short-term observations. We first develop optimal inference techniques for assessing the effects of general treatment patterns. Furthermore, recognizing that many real-world treatments tend to be fine-grained and localized for practical efficiency and operational convenience, we then propose methods to harness this localized structure by sharing information on the non-targeted states. Our new estimator effectively overcomes the variance lower bound for general treatments while matching the more stringent lower bound incorporating the local treatment structure. Furthermore, our estimator can optimally achieve a linear reduction with the number of test arms for a major part of the variance. Finally, we explore scenarios with perfect knowledge of the control arm and design estimators that further improve inference efficiency.

Suggested Citation

Shuze Chen & David Simchi-Levi & Chonghuan Wang, 2024. "Experimenting on Markov Decision Processes with Local Treatments," Papers 2407.19618, arXiv.org, revised Oct 2024.

Handle: RePEc:arx:papers:2407.19618

Download full text from publisher

References listed on IDEAS

Ramesh Johari & Hannah Li & Inessa Liskovich & Gabriel Y. Weintraub, 2022. "Experimental Design in Two-Sided Platforms: An Analysis of Bias," Management Science, INFORMS, vol. 68(10), pages 7069-7089, October.
Duncan I. Simester & Peng Sun & John N. Tsitsiklis, 2006. "Dynamic Catalog Mailing Policies," Management Science, INFORMS, vol. 52(5), pages 683-696, May.
Bruno J.D. Jacobs & Bas Donkers & Dennis Fok, 2016. "Model-Based Purchase Predictions for Large Assortments," Marketing Science, INFORMS, vol. 35(3), pages 389-404, May.
- Jacobs, B.J.D. & Donkers, A.C.D. & Fok, D., 2016. "Model-based Purchase Predictions for Large Assortments," ERIM Report Series Research in Management ERS-2014-007-MKT, Erasmus Research Institute of Management (ERIM), ERIM is the joint research institute of the Rotterdam School of Management, Erasmus University and the Erasmus School of Economics (ESE) at Erasmus University Rotterdam.
Ruoxuan Xiong & Susan Athey & Mohsen Bayati & Guido Imbens, 2024. "Optimal Experimental Design for Staggered Rollouts," Management Science, INFORMS, vol. 70(8), pages 5317-5336, August.
- Ruoxuan Xiong & Susan Athey & Mohsen Bayati & Guido Imbens, 2019. "Optimal Experimental Design for Staggered Rollouts," Papers 1911.03764, arXiv.org, revised Sep 2023.
- Athey, Susan & Imbens, Guido W. & Bayati, Mohsen, 2019. "Optimal Experimental Design for Staggered Rollouts," Research Papers 3837, Stanford University, Graduate School of Business.
Stefan Wager & Kuang Xu, 2021. "Experimenting in Equilibrium," Management Science, INFORMS, vol. 67(11), pages 6694-6715, November.
S. A. Murphy, 2003. "Optimal dynamic treatment regimes," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 65(2), pages 331-355, May.
Yuchen Hu & Stefan Wager, 2022. "Switchback Experiments under Geometric Mixing," Papers 2209.00197, arXiv.org, revised Apr 2024.
Zhan, Ruohan & Hadad, Vitor & Hirshberg, David A. & Athey, Susan, 2021. "Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits," Research Papers 3970, Stanford University, Graduate School of Business.
Iavor Bojinov & David Simchi-Levi & Jinglong Zhao, 2023. "Design and Analysis of Switchback Experiments," Management Science, INFORMS, vol. 69(7), pages 3759-3777, July.
Aurélie Lemmens & Sunil Gupta, 2020. "Managing Churn to Maximize Profits," Marketing Science, INFORMS, vol. 39(5), pages 956-973, September.
Rembrand Koning & Sharique Hasan & Aaron Chatterji, 2022. "Experimentation and Start-up Performance: Evidence from A/B Testing," Management Science, INFORMS, vol. 68(9), pages 6434-6453, September.
Chengchun Shi & Xiaoyu Wang & Shikai Luo & Hongtu Zhu & Jieping Ye & Rui Song, 2023. "Dynamic Causal Effects Evaluation in A/B Testing with a Reinforcement Learning Framework," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 118(543), pages 2059-2071, July.
Jalaj Bhandari & Daniel Russo & Raghav Singal, 2021. "A Finite Time Analysis of Temporal Difference Learning with Linear Function Approximation," Operations Research, INFORMS, vol. 69(3), pages 950-973, May.
Imbens,Guido W. & Rubin,Donald B., 2015. "Causal Inference for Statistics, Social, and Biomedical Sciences," Cambridge Books, Cambridge University Press, number 9780521885881, January.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Xinqi Chen & Xingyu Bai & Zeyu Zheng & Nian Si, 2025. "Bias Analysis of Experiments for Multi-Item Multi-Period Inventory Control Policies," Papers 2501.11996, arXiv.org.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Jinglong Zhao, 2024. "Experimental Design For Causal Inference Through An Optimization Lens," Papers 2408.09607, arXiv.org, revised Aug 2024.
Shan Huang & Chen Wang & Yuan Yuan & Jinglong Zhao & Brocco & Zhang, 2023. "Estimating Effects of Long-Term Treatments," Papers 2308.08152, arXiv.org, revised Dec 2024.
Ke Sun & Linglong Kong & Hongtu Zhu & Chengchun Shi, 2024. "ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments," Papers 2408.05342, arXiv.org, revised Jan 2025.
Ruohan Zhan & Shichao Han & Yuchen Hu & Zhenling Jiang, 2024. "Estimating Treatment Effects under Recommender Interference: A Structured Neural Networks Approach," Papers 2406.14380, arXiv.org, revised Jul 2024.
Luofeng Liao & Christian Kroer, 2024. "Statistical Inference and A/B Testing in Fisher Markets and Paced Auctions," Papers 2406.15522, arXiv.org, revised Mar 2025.
Han, Kevin & Basse, Guillaume & Bojinov, Iavor, 2024. "Population interference in panel experiments," Journal of Econometrics, Elsevier, vol. 238(1).
Ozan Candogan & Chen Chen & Rad Niazadeh, 2024. "Correlated Cluster-Based Randomized Experiments: Robust Variance Minimization," Management Science, INFORMS, vol. 70(6), pages 4069-4086, June.
Xinqi Chen & Xingyu Bai & Zeyu Zheng & Nian Si, 2025. "Bias Analysis of Experiments for Multi-Item Multi-Period Inventory Control Policies," Papers 2501.11996, arXiv.org.
Luofeng Liao & Christian Kroer, 2023. "Statistical Inference and A/B Testing for First-Price Pacing Equilibria," Papers 2301.02276, arXiv.org, revised Jun 2023.
Jinglong Zhao, 2023. "Adaptive Neyman Allocation," Papers 2309.08808, arXiv.org, revised Sep 2023.
Ruoxuan Xiong & Alex Chin & Sean J. Taylor, 2024. "Data-Driven Switchback Experiments: Theoretical Tradeoffs and Empirical Bayes Designs," Papers 2406.06768, arXiv.org.
Evan Munro & David Jones & Jennifer Brennan & Roland Nelet & Vahab Mirrokni & Jean Pouget-Abadie, 2023. "Causal Estimation of User Learning in Personalized Systems," Papers 2306.00485, arXiv.org.
Ruohan Zhan & Zhimei Ren & Susan Athey & Zhengyuan Zhou, 2024. "Policy Learning with Adaptively Collected Data," Management Science, INFORMS, vol. 70(8), pages 5270-5297, August.
- Ruohan Zhan & Zhimei Ren & Susan Athey & Zhengyuan Zhou, 2021. "Policy Learning with Adaptively Collected Data," Papers 2105.02344, arXiv.org, revised Nov 2022.
- Zhan, Ruohan & Ren, Zhimei & Athey, Susan & Zhou, Zhengyuan, 2021. "Policy Learning with Adaptively Collected Data," Research Papers 3963, Stanford University, Graduate School of Business.
Ta-Wei Huang & Eva Ascarza, 2024. "Doing More with Less: Overcoming Ineffective Long-Term Targeting Using Short-Term Signals," Marketing Science, INFORMS, vol. 43(4), pages 863-884, July.
Yusuke Narita, 2018. "Toward an Ethical Experiment," Cowles Foundation Discussion Papers 2127, Cowles Foundation for Research in Economics, Yale University.
Yusuke Narita, 2018. "Experiment-as-Market: Incorporating Welfare into Randomized Controlled Trials," Cowles Foundation Discussion Papers 2127r, Cowles Foundation for Research in Economics, Yale University, revised May 2019.
- Yusuke Narita, 2019. "Experiment-as-Market: Incorporating Welfare into Randomized Controlled Trials," Working Papers 2019-025, Human Capital and Economic Opportunity Working Group.
Luo, Shikai & Yang, Ying & Shi, Chengchun & Yao, Fang & Ye, Jieping & Zhu, Hongtu, 2024. "Policy evaluation for temporal and/or spatial dependent experiments," LSE Research Online Documents on Economics 122741, London School of Economics and Political Science, LSE Library.
Athey, Susan & Imbens, Guido W., 2022. "Design-based analysis in Difference-In-Differences settings with staggered adoption," Journal of Econometrics, Elsevier, vol. 226(1), pages 62-79.
- Susan Athey & Guido Imbens, 2018. "Design-based Analysis in Difference-In-Differences Settings with Staggered Adoption," Papers 1808.05293, arXiv.org, revised Sep 2018.
- Athey, Susan & Imbens, Guido W., 2018. "Design-based Analysis in Difference-In-Differences Settings with Staggered Adoption," Research Papers 3712, Stanford University, Graduate School of Business.
- Susan Athey & Guido W. Imbens, 2018. "Design-based Analysis in Difference-In-Differences Settings with Staggered Adoption," NBER Working Papers 24963, National Bureau of Economic Research, Inc.
Vasilis Syrgkanis & Ruohan Zhan, 2023. "Post Reinforcement Learning Inference," Papers 2302.08854, arXiv.org, revised May 2024.
Valendin, Jan & Reutterer, Thomas & Platzer, Michael & Kalcher, Klaudius, 2022. "Customer base analysis with recurrent neural networks," International Journal of Research in Marketing, Elsevier, vol. 39(4), pages 988-1018.

More about this item

NEP fields

This paper has been announced in the following NEP Reports:

NEP-ECM-2024-08-26 (Econometrics)
NEP-EXP-2024-08-26 (Experimental Economics)

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2407.19618. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: http://arxiv.org/ .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Experimenting on Markov Decision Processes with Local Treatments

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

NEP fields

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data