1.西安交通大学电信学部,710049,西安
2.河北新天科创新能源技术有限公司,075000,河北张家口
收稿:2026-03-30,
修回:2026-07-17,
录用:2026-07-27,
移动端阅览
张哲源, 冯文晖, 连峰, 等. 采用多智能体强化学习的多传感器协同控制策略[J/OL]. 西安交通大学学报, 2026.
ZHANG Zheyuan, FENG Wenhui, LIAN Feng, et al. Multi-Sensor Cooperative Control Strategies Using Multi-Agent Reinforcement Learning[J/OL]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY, 2026.
针对复杂对抗环境下多传感器系统在任务执行与威胁规避过程中的协同决策问题,提出一种采用多智能体强化学习的异构传感器协同控制策略。首先,构建包含主传感器、次传感器、敌方导弹和固定监视目标的二维对抗场景,建立传感器机动、辐射控制、目标探测以及导弹锁定与命中机制模型,并将多传感器协同管理问题建模为多智能体部分可观测马尔可夫决策过程。在集中训练、分散执行框架下,采用多智能体近端策略优化(MAPPO)算法实现协同决策,其中主传感器通过联合优化机动、功率与波束指向完成穿越与监视任务,次传感器通过协同机动与辐射诱导分担主传感器威胁风险。针对强对抗和稀疏反馈条件,设计兼顾任务收益与生存风险的奖励函数,并采用三阶段课程学习策略提升策略训练能力。在包含1个主传感器、2个次传感器和3枚导弹的仿真场景中进行50回合对比评估,结果表明:在覆盖阈值为0.75时,所提方法任务成功率达到0.82,高于威胁感知启发式策略(0.66)、独立近端策略优化策略(0.68)和随机策略(0);同时主传感器存活率达到0.82,平均覆盖率达到0.835。结果表明,所提方法能够有效协调监视收益与辐射暴露风险,提高复杂对抗环境下多传感器系统的任务完成能力与生存能力。
To address the collaborative decision-making problem of multi-sensor systems during mission execution and threat avoidance in complex adversarial environments
this paper proposes a heterogeneous sensor cooperative control policy using multi-agent reinforcement learning. First
a two-dimensional adversarial scenario is constructed involving primary sensors
secondary sensors
hostile missiles
and fixed surveillance targets
and the models for sensor maneuvering
emission control
target detection
as well as missile lock-on and hit mechanisms are established. The multi-sensor cooperative management problem is then formulated as a partially observable Markov decision process within a multi-agent setting. Under the framework of centralized training with decentralized execution
the multi-agent proximal policy optimization (MAPPO) algorithm is adopted to realize cooperative decision-making
where the primary sensor jointly optimizes maneuvering
power
and beam steering to accomplish traversal and surveillance tasks
while secondary sensors share the threat risk from the primary sensor through cooperative maneuvering and radiation induction. To cope with strong adversarial conditions and sparse feedback
a reward function is designed that balances mission payoff and survival risk
and a three-stage curriculum learning strategy is employed to enhance policy training performance. Comparative evaluations are conducted over 50 episodes in a simulation scenario comprising one primary sensor
two secondary sensors
and three missiles. The results show that
with a coverage threshold of 0.75
the proposed method achieves a mission success rate of 0.82
outperforming the threat-aware heuristic policy (0.66)
the independent proximal policy optimization policy (0.68)
and the random policy (0). Meanwhile
the primary sensor survival rate reaches 0.82
and the average coverage rate attains 0.835. These results demonstrate that the proposed method can effectively balance surveillance gains against radiation exposure risks
thereby improving both mission completion capability and survivability of multi-sensor systems in complex adversarial environments.
Bier S G , Rothman P L , Manske R A . Intelligent sensor management for beyond visual range air-to-air combat [C ] // Proceedings of the IEEE 1988 National Aerospace and Electronics Conference . IEEE , 1988 : 264 - 269 .
Ng G W , Ng K H . Sensor management–what, why and how [J ] . Information Fusion , 2000 , 1 ( 2 ): 67 - 75 .
Bar-Shalom Y , Li X R . Multitarget-multisensor tracking: principles and techniques [M ] . Storrs, CT : YBS Publishing , 1995 .
Hero A O , Castanon D A , Cochran D , et al . Foundations and applications of sensor management [M ] . New York : Springer , 2007 .
Zuo Lei , Hu Juan , Sun Hao , et al . Resource allocation for target tracking in multiple radar architectures over lossy networks [J ] . Signal Processing , 2023 , 208 : 108973 .
Hoang H G , Vo B T . Sensor management for multi-target tracking via multi-Bernoulli filtering [J ] . Automatica , 2014 , 50 ( 4 ): 1135 - 1142 .
Jiang Meng , Yi Wei , Kong Lingjiang . Multi-sensor control for multi-target tracking using Cauchy-Schwarz divergence [C ] // 2016 19th International Conference on Information Fusion (FUSION) . IEEE , 2016 : 2059 - 2066 .
Wang Ping , Ma Liang , Xue Kai . Multitarget tracking in sensor networks via efficient information-theoretic sensor selection [J ] . International Journal of Advanced Robotic Systems , 2017 , 14 ( 5 ): 1729881417728466 .
陈辉 , 韩崇昭 . 机动多目标跟踪中的传感器控制策略的研究 [J ] . 自动化学报 , 2016 , 42 ( 4 ): 512 - 523 .
Chen Hui , Han Chongzhao . Research on sensor control strategy for maneuvering multi-target tracking [J ] . Acta Automatica Sinica , 2016 , 42 ( 4 ): 512 - 523 .
Castanon D A . Approximate dynamic programming for sensor management [C ] // Proceedings of the 36th IEEE Conference on Decision and Control . IEEE , 1997 , 2 : 1202 - 1207 .
Krishnamurthy V , Djonin D V . Optimal threshold policies for multivariate POMDPs in radar resource management [J ] . IEEE Transactions on Signal Processing , 2009 , 57 ( 10 ): 3954 - 3969 .
Schöpe M I , Driessen H , Yarovoy A . A constrained POMDP formulation and algorithmic solution for radar resource management in multi-target tracking [J ] . Journal of Advances in Information Fusion , 2021 , 16 ( 1 ): 31 .
王增福 , 杨广宇 , 金术玲 . 考虑综合性能最优的非短视快速天基雷达多目标跟踪资源调度算法 [J ] . 雷达学报 , 2023 , 13 ( 1 ): 253 - 269 .
Wang Zengfu , Yang Guangyu , Jin Shuling . Non-myopic fast space-borne radar resource scheduling algorithm considering comprehensive optimal performance [J ] . Journal of Radars , 2023 , 13 ( 1 ): 253 - 269 .
徐公国 , 单甘霖 , 段修生 . 采用马氏决策过程和后验克拉美罗下界的多被动式移动传感器长期调度方法 [J ] . 西安交通大学学报 , 2019 , 53 ( 6 ): 125 - 133 .
Xu Gongguo , Shan Ganlin , Duan Xiusheng . Long-term scheduling method of multiple passive mobile sensors based on Markov decision process and posterior Cramér-Rao lower bound [J ] . Journal of Xi’an Jiaotong University , 2019 , 53 ( 6 ): 125 - 133 .
Mnih V , Kavukcuoglu K , Silver D , et al . Human-level control through deep reinforcement learning [J ] . Nature , 2015 , 518 ( 7540 ): 529 - 533 .
Charlish A , Hoffmann F , Klemm R , et al . Cognitive radar management [M ] // Novel Radar Techniques and Applications . Routledge, 2017 , 2 ( 3 ): 157 - 193 .
Ravier R , Garagić D , Peskoe J , et al . Online reinforcement learning for autonomous sensor control [C ] // 2023 IEEE Aerospace Conference . IEEE , 2023 : 1 - 10 .
张虹芸 , 陈辉 , 张文旭 . 扩展目标跟踪中基于深度强化学习的传感器管理方法 [J ] . 自动化学报 , 2024 , 50 ( 7 ): 1417 - 1431 .
Zhang Hongyun , Chen Hui , Zhang Wenxu . Sensor management method based on deep reinforcement learning for extended target tracking [J ] . Acta Automatica Sinica , 2024 , 50 ( 7 ): 1417 - 1431 .
闫实 , 贺静 , 王跃东 , 等 . 基于强化学习的多机协同传感器管理 [J ] . 系统工程与电子技术 , 2020 , 42 ( 8 ): 1726 - 1733 .
Yan Shi , He Jing , Wang Yuedong , et al . Multi-aircraft cooperative sensor management based on reinforcement learning [J ] . Systems Engineering and Electronics , 2020 , 42 ( 8 ): 1726 - 1733 .
Zhang Kaiqing , Yang Zhuoran , Başar T . Multi-agent reinforcement learning: a selective overview of theories and algorithms [J ] . Handbook of Reinforcement Learning and Control , 2021 : 321 - 384 .
Lowe R , Wu Yi , Tamar A , et al . Multi-agent actor-critic for mixed cooperative-competitive environments [J ] . Advances in Neural Information Processing Systems , 2017 , 30 : 6382 - 6393 .
Rashid T , Samvelyan M , De Witt C S , et al . Monotonic value function factorisation for deep multi-agent reinforcement learning [J ] . Journal of Machine Learning Research , 2020 , 21 ( 178 ): 1 - 51 .
Yu Chao , Velu A , Vinitsky E , et al . The surprising effectiveness of ppo in cooperative multi-agent games [J ] . Advances in Neural Information Processing Systems , 2022 , 35 : 24611 - 24624 .
Lee J , Niyato D , Guan Yongliang , et al . Learning to schedule joint radar-communication with deep multi-agent reinforcement learning [J ] . IEEE Transactions on Vehicular Technology , 2021 , 71 ( 1 ): 406 - 422 .
Li Xiaoyang , Wang Teng , Wang Yongkun , et al . Intelligent decision-making algorithm for multi-UAV radar cooperative guided search task based on multi-agent reinforcement learning [J ] . IEEE Internet of Things Journal , 2025 , 12 ( 15 ): 31042 - 31063 .
Beatty E , Dong E . Multi-agent reinforcement learning for UAV sensor management [C ] // Open Architecture/Open Business Model Net-Centric Systems and Defense Transformation 2023 . SPIE , 2023 , 12544 : 102 - 111 .
Yu Yue , Liu Mei . Deep reinforcement learning-based multi-sensor control for labeled multi-Bernoulli filtering [J ] . IEEE Transactions on Aerospace and Electronic Systems , 2025 , 61 ( 5 ): 13548 - 13564 .
Cai Bingchen , Li Haoran , Zhang Naimin , et al . A cooperative jamming decision-making method based on multi-agent reinforcement learning [J ] . Autonomous Intelligent Systems , 2025 , 5 ( 1 ): 3 .
Feng Cheng , Fu Xiongjun , Wang Ziyi , et al . An optimization method for collaborative radar antijamming based on multi-agent reinforcement learning [J ] . Remote Sensing , 2023 , 15 ( 11 ): 2893 .
Zhang Wenxu , Zhao Tong , Zhao Zhongkai , et al . An intelligent strategy decision method for collaborative jamming based on hierarchical multi-agent reinforcement learning [J ] . IEEE Transactions on Cognitive Communications and Networking , 2024 , 10 ( 4 ): 1467 - 1480 .
黄洁瑜 , 谢军伟 , 杨子晴 , 等 . 雷达资源管理技术发展研究综述 [J ] . 现代雷达 , 2025 , 47 ( 3 ): 1 - 13 .
Huang Jieyu , Xie Junwei , Yang Ziqing , et al . Review on development of radar resource management technology [J ] . Modern Radar , 2025 , 47 ( 3 ): 1 - 13 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621