Computing Optimal Policies for Markovian Decision Processes Using Simulation

Apostolos N. Burnetas; Michael N. Katehakis

doi:10.1017/S0269964800004034

Computing Optimal Policies for Markovian Decision Processes Using Simulation

Published online by Cambridge University Press: 27 July 2009

Apostolos N. Burnetas and

Michael N. Katehakis

Show author details

Apostolos N. Burnetas: Affiliation:
Weatherhead School of Management, Case Western Reserve University, Cleveland, Ohio 44106
Michael N. Katehakis: Affiliation:
Graduate School of Management and RUTCOR Rutgers University, 92 New Street, Newark, New Jersey 07102-1895

Article contents

Abstract
References

Get access

Rights & Permissions

Abstract

A simulation method is developed for computing average reward optimal policies, for a finite state and action Markovian decision process. It is shown that the method is consistent; i.e., it produces solutions arbitrarily close to the optimal. Various types of estimation errors and confidence bounds are examined. Finally, it is shown that the probability distribution of the number of simulation cycles required to compute an e-optimal policy satisfies a large deviations property.

Information

Type: Research Article
Information: Probability in the Engineering and Informational Sciences , Volume 9 , Issue 4 , October 1995 , pp. 525 - 537

DOI: https://doi.org/10.1017/S0269964800004034 [Opens in a new window]
Copyright: Copyright © Cambridge University Press 1995

Access options

Get access to the full version of this content by using one of the access options below. (Log in options will check for institutional or personal access. Content may require purchase if you do not have access.)

Article purchase

Temporarily unavailable

References

1.Agrawal, R., Teneketzis, D., & Anantharam, V. (1989). Asymptotically efficient adaptive allocation schemes for controlled Markov chains: Finite parameter space. IEEE Transactions on Automated Control 34: 1249–1259.CrossRef Google Scholar

2.Burnetas, A.N. & Katehakis, M.N. (1992). On power one estimation from simulation of finite Markov chains. Technical Report 711-0003, Rutgers University, New Brunswick, NJ.Google Scholar

3.Burnetas, A.N. & Katehakis, M.N. (1994). Optimal adaptive policies for dynamic programming. Technical Report, Rutgers University, New Brunswick, NJ.Google Scholar

4.Crane, M.A. & Iglehart, D.L. (1974). Simulating stable stochastic systems, i: General multiserver queues. Journal of the Association for Computing Machines 21: 103–113.CrossRef Google Scholar

5.Crane, M.A. & Iglehart, D.L. (1974). Simulating stable stochastic systems, ii: Markov chains. Journal of the Association for Computing Machines 21: 114–123.CrossRef Google Scholar

6.Dembo, A. & Zeitouni, O. (1993). Large deviations techniques and applications. Jones and Bartlett.Google Scholar

7.Derman, C. (1970). Finite state Markovian decision processes. New York: Academic Press.Google Scholar

8.Ellis, R.S. (1985). Entropy, large deviations and statistical mechanics. New York: Springer-Verlag.CrossRef Google Scholar

9.Federgruen, A. & Schweitzer, P. (1981). Nonstationary Markov decision problems with converging parameters. Journal of Optimization Theory Applications 34: 207–241.CrossRef Google Scholar

10.Hernández-Lerma, O. (1989). Adaptive Markov control processes. New York: Springer-Verlag.CrossRef Google Scholar

11.Kumar, P.R. (1985). A survey of some results in stochastic adaptive control. SIAM Journal on Control and Optimization 23: 329–380.CrossRef Google Scholar

12.Ross, S.M. (1983). Introduction to stochastic dynamic programming. New York: Academic Press.Google Scholar

13.Ross, S.M. & Schechner, Z. (1985). Using simulation to estimate first passage distribution. Management Science 31(2): 224–234.CrossRef Google Scholar

14.Thomas, L.C., Harley, R. & Lavercombe, A.C. (1983). Computational comparisons of value iteration algorithms for discounted Markov decision processes. Operations Research Letters 2: 72–76.CrossRef Google Scholar

15.Van-Dijk, N.M. & Puterman, M.L. (1988). Perturbation theory for Markov reward processes with applications to queueing systems. Advances in Applied Probability 20: 79–98.CrossRef Google Scholar

Article contents

Computing Optimal Policies for Markovian Decision Processes Using Simulation

Abstract

Information

Access options

Article purchase

Temporarily unavailable

References

Save article to Kindle

Save article to Dropbox

Save article to Google Drive

Reply to: Submit a response

Your details

You have entered the maximum number of contributors

Conflicting interests