A Q-learning algorithm based on predictive state representations is proposed for solving the problem of planning under uncertainty. The predictive state representations is combined with the Q-learning algorithm. The prediction vector of predictive state representations is used as the state representation of Q-learning algorithms
so that the obtained states have the Markov properties and satisfy the requirement of reinforcement learning tasks. Then the Q-learning algorithm is used to find the optimal policy and the problem of planning under uncertainty is solved. Simulation results show that with our algorithm
the number of episodes needed in finding the near-optimal policy of an agent is approximately the same as that of the world states being assumed to be known.
关键词
Keywords
references
KAELBLING L P, LITTMAN M L, CASSANDRA A R. Planning and acting in partially observable stochastic domains [J]. Artificial Intelligence, 1998, 101(1/2): 99-134.
LITTLEMAN M L, SUTTON R S, SINGH S. Predictive representation of state [M]∥Advances in Neural Information Processing Systems 14. Cambridge, MA, USA: MIT Press, 2002: 1555-1561.
SINGH S, JAMES M R, RUDARY M R. Predictive state representations: a new theory for modeling dynamical systems [C]∥Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence. Alberta, Canada: AUAI Press: 512-519.
MCCALLUM A K. Reinforcement learning with selective perception and hidden state [D].University of Rochester. Department of Computer Science, 1995.
JAMES M R, SINGH S. Learning and discovery of predictive state representations in dynamical systems with reset [C]∥Proceedings of the 21st International Conference on Machine Learning. New York, USA ACM,2004: 417-424.
WOLFE B, JAMES M R, SINGH S. Learning predictive state representations in dynamical systems without reset [C]∥Proceeding of the 22nd International Conference on Machine Learning. New York, USA ACM, 2005: 985-992.
SUTTON R S, BRATO A G. Reinforcement learning: an introduction[M]. Cambridge, MA, USA: MIT Press, 1998.