西安交通大学系统工程研究所,西安,710049
网络首发:2008-12-10,
纸质出版:2008
移动端阅览
刘云龙, 李人厚, 刘建书. 基于预测状态表示的Q学习算法[J]. 西安交通大学学报, 2008,42(12):1472-1475+1485.
刘云龙, 李人厚, 刘建书. Q-Learning Algorithm Based on Predictive State Representations[J]. 2008, 42(12): 1472-1475+1485.
针对不确定环境的规划问题
提出了基于预测状态表示的Q学习算法.将预测状态表示方法与Q学习算法结合
用预测状态表示的预测向量作为Q学习算法的状态表示
使得到的状态具有马尔可夫特性
满足强化学习任务的要求
进而用Q学习算法学习智能体的最优策略
可解决不确定环境下的规划问题.仿真结果表明
在发现智能体的最优近似策略时
算法需要的学习周期数与假定环境状态已知情况下需要的学习周期数大致相同.
A Q-learning algorithm based on predictive state representations is proposed for solving the problem of planning under uncertainty. The predictive state representations is combined with the Q-learning algorithm. The prediction vector of predictive state representations is used as the state representation of Q-learning algorithms
so that the obtained states have the Markov properties and satisfy the requirement of reinforcement learning tasks. Then the Q-learning algorithm is used to find the optimal policy and the problem of planning under uncertainty is solved. Simulation results show that with our algorithm
the number of episodes needed in finding the near-optimal policy of an agent is approximately the same as that of the world states being assumed to be known.
KAELBLING L P, LITTMAN M L, CASSANDRA A R. Planning and acting in partially observable stochastic domains [J]. Artificial Intelligence, 1998, 101(1/2): 99-134.
LITTLEMAN M L, SUTTON R S, SINGH S. Predictive representation of state [M]∥Advances in Neural Information Processing Systems 14. Cambridge, MA, USA: MIT Press, 2002: 1555-1561.
SINGH S, JAMES M R, RUDARY M R. Predictive state representations: a new theory for modeling dynamical systems [C]∥Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence. Alberta, Canada: AUAI Press: 512-519.
MCCALLUM A K. Reinforcement learning with selective perception and hidden state [D].University of Rochester. Department of Computer Science, 1995.
JAMES M R, SINGH S. Learning and discovery of predictive state representations in dynamical systems with reset [C]∥Proceedings of the 21st International Conference on Machine Learning. New York, USA ACM,2004: 417-424.
WOLFE B, JAMES M R, SINGH S. Learning predictive state representations in dynamical systems without reset [C]∥Proceeding of the 22nd International Conference on Machine Learning. New York, USA ACM, 2005: 985-992.
SUTTON R S, BRATO A G. Reinforcement learning: an introduction[M]. Cambridge, MA, USA: MIT Press, 1998.
0
浏览量
5
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621