

浏览全部资源
扫码关注微信
1. 西安交通大学热流科学与工程教育部重点实验室,西安,710049
2. 西安热工研究院有限公司,西安,710054
3. 东南大学能源热转换及其过程测控教育部重点实验室,南京,210096
Online First:10 August 2022,
Published:2022
移动端阅览
ZHOU Dongyang, CAO Jun, BI Shengshan, et al. Reinforcement-Learning-Based Performance Optimal Control Framework and Its Application in Operation Optimization of High Pressure Feedwater Heaters[J]. 2022, 56(8): 32-42.
ZHOU Dongyang, CAO Jun, BI Shengshan, et al. Reinforcement-Learning-Based Performance Optimal Control Framework and Its Application in Operation Optimization of High Pressure Feedwater Heaters[J]. 2022, 56(8): 32-42. DOI: 10.7652/xjtuxb202208004.
针对现阶段火电机组运行工况频繁波动的情况
为了解决复杂动态过程难以辨识、控制器设定点无法确定的问题
提出了一种基于历史运行数据与强化学习算法的性能最优控制框架。在现有控制器的输出上叠加少量随机噪声
采用均匀化网格算法构建并维护包含典型工况的数据缓冲区
采用基于粒子群优化的连续批量Q学习算法离线求解性能最优控制策略函数。以高压给水加热器控制任务为研究对象
得到了一种无需系统辨识也无需确定设定点即可保持变工况控制品质与换热性能的控制器求解方法。为了验证所提框架的通用性
利用某600 MW机组高压加热器的仿真模型对水位控制过程进行了分析。结果表明
基于强化学习的性能最优控制框架不需要建立系统模型
可以直接利用历史运行数据求解以累积性能最优为目标的控制策略函数
不仅在动态过程中可以达到较好的控制品质
稳态下也能使系统维持在性能较优的状态
相当于同时实现了设定值优化与设定点跟踪控制。
In view of the frequent fluctuations in the operating conditions of thermal power units
to solve the problems that the complex dynamic process is difficult to be identified and that the controller setpoint cannot be determined
this paper proposes a performance optimal control framework based on historical operating data and reinforcement learning algorithms. Firstly
a small amount of random noise is superimposed on the output of the existing controller; then a data buffer containing typical operating conditions is built and maintained with the homogenization grid algorithm; and finally the performance optimal control policy function is solved offline with batch Q-learning algorithm based on particle swarm optimization. Focusing on the research into the control task of high-pressure feedwater heaters
this paper proposes a controller solving method that can maintain the control quality and heat transfer performance under variable operating conditions without system identification or setpoint determination. In order to verify the versatility of the performance optimal control framework
the water level control process was analyzed by using the simulation model of a high-pressure heater of a 600 MW unit. The results show that the reinforcement-learning-based performance optimal control framework works without the need of establishing a system model
and can directly solve the control policy function using historical data for the optimal cumulative performance
which not only realizes better control in the dynamic process
but also keeps the system in its optimal performance in a steady state. This method achieves setpoint optimization and setpoint tracking control at the same time.
LEWIS F L, VRABIE D L, SYRMOS V L. Optimal control [M]. Hoboken, NJ, USA: Wiley, 2012.
WHITE D J. Dynamic programming, Markov chains, and the method of successive approximations [J]. Journal of Mathematical Analysis and Applications, 1963, 6(3): 373-376.
ÅSTRÖM K J. Adaptive control [M]. [S.l.]: Courier Corporation, 2013.
DOYLE J. Robust and optimal control [C]∥Proceedings of 35th IEEE Conference on Decision and Control. Piscataway, NJ, USA: IEEE, 1996: 1595-1598.
李书臣, 徐心和, 李平. 预测控制最新算法综述 [J]. 系统仿真学报, 2004, 16(6): 1314-1319, 1349.
LI Shuchen, XU Xinhe, LI Ping. An overview of new algorithms in predictive control [J]. Journal of System Simulation, 2004, 16(6): 1314-1319, 1349.
LEWIS F L, VRABIE D. Reinforcement learning and adaptive dynamic programming for feedback control [J]. IEEE Circuits and Systems Magazine, 2009, 9(3): 32-50.
HOSSIENALIPOUR S M, KARBALAEE M S, FATHIANNASAB H. Development of a model to evaluate the water level impact on drain cooling in horizontal high pressure feedwater heaters [J]. Applied Thermal Engineering, 2017, 110: 590-600.
XU Jianqun, YANG Tao, SUN Youyuan, et al. Research on varying condition characteristic of feedwater heater considering liquid level [J]. Applied Thermal Engineering, 2014, 67(1/2): 179-189.
ANTAR M A, ZUBAIR S M. The impact of fouling on performance evaluation of multi-zone feedwater heaters [J]. Applied Thermal Engineering, 2007, 27(14/15): 2505-2513.
CHAI Tianyou, QIN S J, WANG Hong. Optimal operational control for complex industrial processes [J]. Annual Reviews in Control, 2014, 38(1): 81-92.
YIN Shen, LUO Hao, DING S X. Real-time implementation of fault-tolerant control systems with performance optimization [J]. IEEE Transactions on Industrial Electronics, 2014, 61(5): 2402-2411.
LU Xinglong, KIUMARSI B, CHAI Tianyou, et al. Data-driven optimal control of operational indices for a class of industrial processes [J]. IET Control Theory Applications, 2016, 10(12): 1348-1356.
DAI Wei, CHAI Tianyou, YANG S X. Data-driven optimization control for safety operation of hematite grinding process [J]. IEEE Transactions on Industrial Electronics, 2015, 62(5): 2930-2941.
WANG Ding, HE Haibo, MU Chaoxu, et al. Intelligent critic control with disturbance attenuation for affine dynamics including an application to a microgrid system [J]. IEEE Transactions on Industrial Electronics, 2017, 64(6): 4935-4944.
WANG Ding, HE Haibo, LIU Derong. Improving the critic learning for event-based nonlinear H∞ control design [J]. IEEE Transactions on Cybernetics, 2017, 47(10): 3417-3428.
THRUN S. Learning to play the game of chess [M]∥Advances in Neural Information Processing Systems. Cambridge, MA, USA: The MIT Press, 1995: 1069-1076.
SAMUEL A L. Some studies in machine learning using the game of checkers [J]. IBM Journal of Research and Development, 2000, 44(1/2): 206-226.
吴晓军, 张成, 原盛, 等. 基于强化学习的云资源混合式弹性伸缩算法 [J]. 西安交通大学学报, 2022, 56(1): 142-150.
WU Xiaojun, ZHANG Cheng, YUAN Sheng, et al. Blended elastic scaling method for cloud resources following reinforcement learning [J]. Journal of Xi'an Jiaotong University, 2022, 56(1): 142-150.
MNIH V, KAVUKCUOGLU K, SILVER D, et al. Human-level control through deep reinforcement learning [J]. Nature, 2015, 518(7540): 529-533.
SILVER D, SCHRITTWIESER J, SIMONYAN K, et al. Mastering the game of Go without human knowledge [J]. Nature, 2017, 550(7676): 354-359.
KHAN S G, HERRMANN G, LEWIS F L, et al. Reinforcement learning and optimal adaptive control: an overview and implementation examples [J]. Annual Reviews in Control, 2012, 36(1): 42-59.
SUTTON R S, BARTO A G. Reinforcement learning: an introduction [M]. 2nd ed. Cambridge, MA, USA: The MIT Press, 2018.
张铁, 廖才磊, 邹焱飚, 等. 采用强化学习的多轴运动系统时间最优轨迹优化 [J]. 西安交通大学学报, 2021, 55(7): 33-40.
ZHANG Tie, LIAO Cailei, ZOU Yanbiao, et al. Time-optimal trajectory optimization of multi-axis motion system by reinforcement learning [J]. Journal of Xi'an Jiaotong University, 2021, 55(7): 33-40.
ERNST D, GEURTS P, WEHENKEL L. Tree-based batch mode reinforcement learning [J]. The Journal of Machine Learning Research, 2005, 6: 503-556.
KAMPEN E V, CHU Q P, MULDER J A. Online adaptive critic flight control using approximated plant dynamics [C]∥2006 International Conference on Machine Learning and Cybernetics. Piscataway, NJ, USA: IEEE, 2006: 256-261.
LIU Wei, TAN Ying, QIU Qinru. Enhanced Q-learning algorithm for dynamic power management with performance constraint [C]∥2010 Design, Automation Test in Europe Conference Exhibition(DATE 2010). Piscataway, NJ, USA: IEEE, 2010: 602-605.
STINGU P E, LEWIS F L. Adaptive dynamic programming applied to a 6DoF quadrotor [M]∥Computational Modeling and Simulation of Intellect: Current State and Future Perspectives. Hershey, PA, USA: IGI Global, 2011: 102-130.
KOBER J, BAGNELL J A, PETERS J. Reinforcement learning in robotics: a survey [J]. The International Journal of Robotics Research, 2013, 32(11): 1238-1274.
JIANG Yi, FAN Jialu, CHAI Tianyou, et al. Data-driven flotation industrial process operational optimal control based on reinforcement learning [J]. IEEE Transactions on Industrial Informatics, 2018, 14(5): 1974-1989.
赵则飞, 赵轶韬, 赵四海, 等. 600 MW汽轮机高压加热器水位运行优化 [J]. 内蒙古电力技术, 2017, 35(3): 54-56, 60.
ZHAO Zefei, ZHAO Yitao, ZHAO Sihai, et al. Water level operation optimization of high-pressure heater in 600 MW unit [J]. Inner Mongolia Electric Power, 2017, 35(3): 54-56, 60.
杨维, 李歧强. 粒子群优化算法综述 [J]. 中国工程科学, 2004, 6(5): 87-94.
YANG Wei, LI Qiqiang. Survey on particle swarm optimization algorithm [J]. Engineering Science, 2004, 6(5): 87-94.
王珊, 刘明, 严俊杰. 采用粒子群算法的热电厂热电负荷分配优化 [J]. 西安交通大学学报, 2019, 53(9): 159-166.
WANG Shan, LIU Ming, YAN Junjie. Optimizing heat-power load distribution of thermal power plants based on particle swarm algorithm [J]. Journal of Xi'an Jiaotong University, 2019, 53(9): 159-166.
RIEDMILLER M. Neural fitted Q iteration: first experiences with a data efficient neural reinforcement learning method [C]∥16th European Conference on Machine Learning. Cham, Germany: Springer, 2005: 317-328.
STARKLOFF R, ALOBAID F, KARNER K, et al. Development and validation of a dynamic simulation model for a large coal-fired power plant [J]. Applied Thermal Engineering, 2015, 91: 496-506.
0
Views
15
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621