

浏览全部资源
扫码关注微信
1.长安大学电子与控制工程学院, 710064,西安
2.西安交通大学自动化科学与工程学院, 710049,西安
Received:24 February 2025,
Online First:27 April 2025,
Published:10 September 2025
移动端阅览
MA Yu, AN Dou, LIN Xixiang, et al. Hierarchical Twin-Delayed Policy Gradient Reinforcement Learning for Intelligent Cooperative Control of Aircraft[J]. Journal of Xi’an Jiaotong University, 2025, 59(9): 88-98.
MA Yu, AN Dou, LIN Xixiang, et al. Hierarchical Twin-Delayed Policy Gradient Reinforcement Learning for Intelligent Cooperative Control of Aircraft[J]. Journal of Xi’an Jiaotong University, 2025, 59(9): 88-98. DOI: 10.7652/xjtuxb202509009.
针对多飞行器智能协同控制中因规模大、环境复杂及资源受限导致的建模与协同难题,以提高决策算法效率为目标,构建了多智能体分层决策架构,提出了智能协同控制方法。首先,将飞行器作为智能体构建协同控制模型;其次,采用部分可观测马尔可夫决策过程模型解决观测信息不全问题;然后,针对博弈环境多变和学习成本问题,提出基于集中训练分布执行的分层双时延策略梯度强化学习方法,融合有模型(model-based)与无模型(model-free)机制高效利用现有博弈环境的演化模型;最后,在分层智能决策框架下,进行典型多飞行器博弈及千次多场景的仿真验证。结果表明,新方法有效解决多飞行器协同控制问题,相较于多智能体强化学习算法MAPPO和QMIX,训练时间分别减少了51.03%和79.03%,算法效率(累积回报)分别提升了37.51%和58.73%,规避机动成功率分别提高了17.63%和39.79%。
To address the modeling and coordination challenges in intelligent cooperative control of aircraft caused by large-scale systems
complex environments
and resource constraints
this study proposes an intelligent cooperative control method by establishing a hierarchical multi-agent decision-making architecture with the goal of improving decision-making algorithm efficiency. First
aircraft is modeled as an intelligent agent to establish a cooperative control framework. Second
a partially observable Markov decision process (POMDP) model is employed to handle incomplete observation information. Then
to tackle the issues of dynamic game environments and high learning costs
a hierarchical twin-delayed policy gradient reinforcement learning method based on centralized training with decentralized execution is proposed
which effectively combines model-based and model-free mechanisms to leverage existing game environment evolution models. Finally
under the hierarchical decision-making framework
simulations of typical multi-aircraft game scenarios and thousands of multi-scenario tests are conducted. The results demonstrate that the proposed method successfully resolves multi-aircraft cooperative control problem. Compared to the multi-agent reinforcement learning algorithms MAPPO and QMIX
the training time is reduced by 51.03% and 79.03%
algorithm efficiency (cumulative reward) is improved by 37.51% and 58.73%
and evasion maneuver success rate is increased by 17.63% and 39.79%
respectively.
郑卓 , 路坤锋 , 王昭磊 , 等 . 飞行器集群协同控制技术分析与展望 [J ] . 宇航学报 , 2023 , 44 ( 4 ): 538 - 545 .
ZHENG Zhuo , LU Kunfeng , WANG Zhaolei , et al . Analysis and prospect of vehicle swarm cooperative control technology [J ] . Journal of Astronautics , 2023 , 44 ( 4 ): 538 - 545 .
方峰 , 蔡远利 . 三体对抗中的自适应协同突防策略 [J ] . 西安交通大学学报 , 2017 , 51 ( 4 ): 72 - 78 .
FANG Feng , CAI Yuanli . An adaptive collaborative guidance strategy in three-body engagement [J ] . Journal of Xi'an Jiaotong University , 2017 , 51 ( 4 ): 72 - 78 .
FANG Feng , CAI Yuanli , JABBARI F . 3D optimal defensive guidance strategy with safe distance [J ] . Transactions of the Institute of Measurement and Control , 2019 , 41 ( 15 ): 4285 - 4300 .
鲜勇 , 田海鹏 , 王剑 , 等 . 基于微分对策的导弹智能机动突防研究 [J ] . 飞行力学 , 2014 , 32 ( 1 ): 70 - 73 .
XIAN Yong , TIAN Haipeng , WANG Jian , et al . Research on intelligent maneuver penetration of missile based on differential game theory [J ] . Flight Dynamics , 2014 , 32 ( 1 ): 70 - 73 .
BARDHAN R , GHOSE D . Nonlinear differential games-based impact-angle-constrained guidance law [J ] . Journal of Guidance, Control, and Dynamics , 2015 , 38 ( 3 ): 384 - 402 .
任章 , 郭栋 , 董希旺 , 等 . 飞行器集群协同制导控制方法及应用研究 [J ] . 导航定位与授时 , 2019 , 6 ( 5 ): 1 - 9 .
REN Zhang , GUO Dong , DONG Xiwang , et al . Research on the cooperative guidance and control method and application for aerial vehicle swarm systems [J ] . Navigation Positioning and Timing , 2019 , 6 ( 5 ): 1 - 9 .
ZOU Xiaofei , YANG Ruopeng , YIN Changsheng , et al . Deploying tactical communication node vehicles with AlphaZero algorithm [J ] . IET Communications , 2020 , 14 ( 9 ): 1392 - 1396 .
MNIH V , KAVUKCUOGLU K , SILVER D , et al . Human-level control through deep reinforcement learning [J ] . Nature , 2015 , 518 ( 7540 ): 529 - 533 .
谭浪 , 巩庆海 , 王会霞 . 基于深度强化学习的追逃博弈算法 [J ] . 航天控制 , 2018 , 36 ( 6 ): 3 - 8 .
TAN Lang , GONG Qinghai , WANG Huixia . Pursuit-evasion game algorithm based on deep reinforcement learning [J ] . Aerospace Control , 2018 , 36 ( 6 ): 3 - 8 .
YANG Qiming , ZHANG Jiandong , SHI Guoqing , et al . Maneuver decision of UAV in short-range air combat based on deep reinforcement learning [J ] . IEEE Access , 2020 , 8 : 363 - 378 .
FAN Zihao , XU Yang , KANG Yuhang , et al . Air combat maneuver decision method based on A3C deep reinforcement learning [J ] . Machines , 2022 , 10 ( 11 ): 1033 .
ZHANG Jiandong , YANG Qiming , SHI Guoqing , et al . UAV cooperative air combat maneuver decision based on multi-agent reinforcement learning [J ] . Journal of Systems Engineering and Electronics , 2021 , 32 ( 6 ): 1421 - 1438 .
ZHUANG Xing , LI Dongguang , WANG Yue , et al . Optimization of high-speed fixed-wing UAV penetration strategy based on deep reinforcement learning [J ] . Aerospace Science and Technology , 2024 , 148 : 109089 .
WANG Huan , WANG Jintao . Enhancing multi-UAV air combat decision making via hierarchical reinforcement learning [J ] . Scientific Reports , 2024 , 14 ( 1 ): 4458 .
WU Mingyu , HE Xianjun , QIU Zhiming , et al . Guidance law of interceptors against a high-speed maneuvering target based on deep Q-Network [J ] . Transactions of the Institute of Measurement and Control , 2022 , 44 ( 7 ): 1373 - 1387 .
裴培 , 何绍溟 , 王江 , 等 . 一种深度强化学习制导控制一体化算法 [J ] . 宇航学报 , 2021 , 42 ( 10 ): 1293 - 1304 .
PEI Pei , HE Shaoming , WANG Jiang , et al . Integrated guidance and control for missile using deep reinforcement learning [J ] . Journal of Astronautics , 2021 , 42 ( 10 ): 1293 - 1304 .
CHEN Wenxue , GAO Changsheng , JING Wuxing . Proximal policy optimization guidance algorithm for intercepting near-space maneuvering targets [J ] . Aerospace Science and Technology , 2023 , 132 : 108031 .
GONG Xiaopeng , CHEN Wanchun , CHEN Zhongyuan . Intelligent game strategies in target-missile-defender engagement using curriculum-based deep reinforcement learning [J ] . Aerospace , 2023 , 10 ( 2 ): 133 .
JIANG Liang , NAN Ying , ZHANG Yu , et al . Anti-interception guidance for hypersonic glide vehicle: a deep reinforcement learning approach [J ] . Aerospace , 2022 , 9 ( 8 ): 424 .
GAO Mengjing , YAN Tian , LI Quancheng , et al . Intelligent pursuit-evasion game based on deep reinforcement learning for hypersonic vehicles [J ] . Aerospace , 2023 , 10 ( 1 ): 86 .
GUO Yunhe , JIANG Zijian , HUANG Hanqiao , et al . Intelligent maneuver strategy for a hypersonic pursuit-evasion game based on deep reinforcement learning [J ] . Aerospace , 2023 , 10 ( 9 ): 783 .
YAN Tian , JIANG Zijian , LI Tong , et al . Intelligent maneuver strategy for hypersonic vehicles in three-player pursuit-evasion games via deep reinforcement learning [J ] . Frontiers in Neuroscience , 2024 , 18 : 1362303 .
JIANG Liang , NAN Ying , LI Zhihan . Realizing midcourse penetration with deep reinforcement learning [J ] . IEEE Access , 2021 , 9 : 89812 - 89822 .
南英 , 蒋亮 . 基于深度强化学习的弹道导弹中段突防控制 [J ] . 指挥信息系统与技术 , 2020 , 11 ( 4 ): 1 - 9 .
NAN Ying , JIANG Liang . Midcourse penetration and control of ballistic missile based on deep reinforcement learning [J ] . Command Information System and Technology , 2020 , 11 ( 4 ): 1 - 9 .
WANG Yaokun , ZHAO Kun , GUIRAO J L G , et al . Online intelligent maneuvering penetration methods of missile with respect to unknown intercepting strategies based on reinforcement learning [J ] . Electronic Research Archive , 2022 , 30 ( 12 ): 4366 - 4381 .
QIU Xiaoqi , GAO Changsheng , JING Wuxing . Maneuvering penetration strategies of ballistic missiles based on deep reinforcement learning [J ] . Proceedings of the Institution of Mechanical Engineers: Part G Journal of Aerospace Engineering , 2022 , 236 ( 16 ): 3494 - 3504 .
YAN Mengda , YANG Rennong , ZHANG Ying , et al . A hierarchical reinforcement learning method for missile evasion and guidance [J ] . Scientific Reports , 2022 , 12 ( 1 ): 18888 .
YU Chao , VELU A , VINITSKY E , et al . The surprising effectiveness of PPO in cooperative multi-agent games [C ] // Proceedings of the 36th International Conference on Neural Information Processing Systems . Red Hook, NY, USA : Curran Associates Inc. , 2022 : 24611 - 24624 .
RASHID T , SAMVELYAN M , DE WITT C S , et al . Monotonic value function factorisation for deep multi-agent reinforcement learning [J ] . The Journal of Machine Learning Research , 2020 , 21 ( 1 ): 7234 - 7284 .
ZHAO Yu , ZHOU Ding , BAI Chengchao , et al . Reinforcement learning based spacecraft autonomous evasive maneuvers method against multi-interceptors [C ] // 2020 3rd International Conference on Unmanned Systems (ICUS) . Piscataway, NJ, USA : IEEE , 2020 : 1108 - 1113 .
王建波 , 孙冉 , 刘忠凯 , 等 . 面向储能辅助火电机组一次调频的深度强化学习控制策略 [J ] . 西安交通大学学报 , 2024 , 58 ( 6 ): 186 - 192 .
WANG Jianbo , SUN Ran , LIU Zhongkai , et al . Deep reinforcement learning control strategy for primary frequency regulation of energy storage assisted thermal power units [J ] . Journal of Xi'an Jiaotong University , 2024 , 58 ( 6 ): 186 - 192 .
0
Views
9
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621