1. 长春理工大学机电工程学院,长春,130022
2. 长春理工大学重庆研究院,重庆,401135
3. 西安交通大学信息与通信工程学院,西安,710049
: 2022-04-14。作者简介: 薛珊(1978—),女,教授,博士生导师。基金项目: 吉林省重点科技研发资助项目(20180201058SF)
网络首发:2022-10-10,
纸质出版:2022
移动端阅览
薛珊, 卫立炜, 顾宸瑜, 等. 采用混合域注意力机制的无人机识别方法[J]. 西安交通大学学报, 2022,56(10):141-150.
Drone Identification Method Based on Mixed Domain Attention Mechanism[J]. 2022, 56(10): 141-150.
薛珊, 卫立炜, 顾宸瑜, 等. 采用混合域注意力机制的无人机识别方法[J]. 西安交通大学学报, 2022,56(10):141-150. DOI: 10.7652/xjtuxb202210014.
Drone Identification Method Based on Mixed Domain Attention Mechanism[J]. 2022, 56(10): 141-150. DOI: 10.7652/xjtuxb202210014.
针对在城市公园、广场和大型游乐场等公共环境中
雷达和无线电识别无人机易受到电子干扰、图像识别无人机易受到光线和遮挡物干扰的问题
提出了一种经济便捷、不易受到干扰的运用声音和采用通道空间混合域注意力机制多尺度分组卷积网络(ECSANet)的无人机识别方法。首先
建立民用的9大类无人机声音数据集
提取数据集的对数梅尔谱图及其动态特征; 其次
为了网络参数量少
避免过拟合
设计了基于分组卷积、通道混洗和残差结构的通道混洗多尺度分组卷积网络(MSSGNet); 然后
为了能更多、更有效地提取无人机声音特征
设计了通道空间混合域注意力机制模块(ECSA); 最后
将ECSA模块插入MSSGNet网络构成改进的通道空间混合域注意力机制的多尺度分组卷积网络(ECSANet)
形成新型声音识别无人机的方法。运用设计的ECSANet网络对自建的民用无人机声音数据集和Urbansound8K环境声音数据集进行了声音识别
识别结果表明:与ResNet18、ResNet34、ResNeXt18和MobileNetV2等基准网络相比
MSSGNet网络参数更少
识别准确率更高
达到了95.1%; ECSA模块可以插入多种网络
在不增加很多参数的情况下令网络模型的识别准确率获得提升
在无人机等声音分类任务上具有很好的效果; 与MSSGNet网络相比
改进的ECSANet网络识别准确率能达到95.9%
提高了0.8%
表明了该网络在识别小样本无人机方面的优越性和可行性。
An economical
convenient and undisturbed drone detection method using sound and multiscale group convolution network with attention mechanism in mixed domain of channel space(ECSANet)is proposed in the context of susceptibility to electronic interference in identification of drones by radar and radio
and the interference of light and obstruction in identification of drones by images in public environments such as urban parks
squares and large amusement parks. Firstly
nine kinds of sound dataset of civil drones are established
and their logarithmic Mel spectra and dynamic characteristics are extracted. Secondly
based on packet convolution
channel shuffling and residual structure
a multi-scale group convolution network with channel shuffle(MSSGNet)is designed to reduce the network parameters and avoid over fitting. Then
the efficient channel and spatial attention(ECSA)is designed to extract more and more effective features of drone sounds. Finally
the ECSA is inserted into the MSSGNet to form an improved multiscale group convolution network with attention mechanism in mixed domain of channel space(ECSANet)
offering a new method for sound recognition of drones. The designed ECSANet is used to identify the self-built civil drone sound dataset and environmental sound dataset urbansound8k. The results reveal that when compared with benchmark networks such as ResNet18
ResNet34
ResNeXt18
and MobileNetV2
the MSSGNet has fewer network parameters but a higher identification accuracy(up to 95.1%). The ECSA can be inserted into a variety of networks to improve identification accuracy of network models without introducing too many parameters
and it works well for sound classification tasks like drones. As compared with the MSSGNet
the improved ECSANet has an identification accuracy of 95.9%
an increase of 0.8 percent
demonstrating the superiority and feasibility in identifying a small sample of drones.
罗俊海, 王芝燕. 无人机探测与对抗技术发展及应用综述 [J]. 控制与决策, 2022, 37(3): 530-544.
LUO Junhai, WANG Zhiyan. A review of development and application of UAV detection and counter technology [J]. Control and Decision, 2022, 37(3): 530-544.
陈唯实, 黄毅峰, 卢贤锋. 多传感器融合的无人机探测技术应用综述 [J]. 现代雷达, 2020, 42(6): 15-29.
CHEN Weishi, HUANG Yifeng, LU Xianfeng.Survey on application of multi-sensor fusion in UAV detection technology [J]. Modern Radar, 2020, 42(6): 15-29.
AL-EMADI S, AL-ALI A, AL-ALI A. Audio-based drone detection and identification using deep learning techniques with dataset enhancement through generative adversarial networks [J]. Sensors, 2021, 21(15): 4953.
ISMAIL M A A, BIERIG A. Identifying drone-related security risks by a laser vibrometer-based payload identification system [C]//Proceedings of Laser Radar Technology and Applications. Bellingham, WA, USA: SPIE, 2018: 1063603.
SEO Y, JANG B, IM S. Drone detection using convolutional neural networks with acoustic STFT features [C]//2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance. Piscataway, NJ, USA: IEEE, 2018: 1-6.
CASABIANCA P, ZHANG Yu. Acoustic-based UAV detection using late fusion of deep neural networks [J]. Drones, 2021, 5(3): 54.
KRIZHEVSKY A, SUTSKEVER I, HINTON G. Image net classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60(6): 84-90.
李云红, 梁思程, 贾凯莉, 等. 一种改进的DNN-HMM的语音识别方法 [J]. 应用声学, 2019, 38(3): 371-377.
LI Yunhong, LIANG Sicheng, JIA Kaili, et al. An improved speech recognition method based on DNN-HMM model [J]. Applied Acoustics, 2019, 38(3): 371-377.
DUMITRESCU C, MINEA M, COSTEA I M, et al. Development of an acoustic system for UAV detection [J]. Sensors, 2020, 20(17): 4870.
JIN Shuaipu, WANG Xiufeng, DU Leilei, et al. Evaluation and modeling of automotive transmission whine noise quality based on MFCC and CNN [J]. Applied Acoustics, 2021, 172: 107562.
SHUKLA J, BARREDA-ÁNGELES M, OLIVER J, et al. Feature extraction and selection for emotion recognition from electrodermal activity [J]. IEEE Transactions on Affective Computing, 2021, 12(4): 857-869.
ZHANG Xiangyu, ZHOU Xinyu, LIN Mengxiao, et al. ShuffleNet: an extremely efficient convolutional neural network for mobile devices [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 6848-6856.
HE Kaiming, ZHANG Xiangyu, REN Shaoqing, et al. Deep residual learning for image recognition [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2016: 770-778.
HOWARD A G, ZHU Menglong, CHEN Bo, et al. MobileNets: efficient convolutional neural networks for mobile vision applications [EB/OL].[2022-01-02]. https://arxiv.org/abs/1704.04861.
余浩帅, 汤宝平, 张楷, 等. 小样本下混合自注意力原型网络的风电齿轮箱故障诊断方法 [J]. 中国机械工程, 2021, 32(20): 2475-2481.
YU Haoshuai, TANG Baoping, ZHANG Kai, et al. Fault diagnosis method of wind turbine gearboxes mixed with attention prototype networks under small samples [J]. China Mechanical Engineering, 2021, 32(20): 2475-2481.
郑作武, 邵斯绮, 高晓沨, 等. 基于社交圈层和注意力机制的信息热度预测 [J]. 计算机学报, 2021, 44(5): 921-936.
ZHENG Zuowu, SHAO Siqi, GAO Xiaofeng, et al. Social circle and attention based information popularity prediction [J]. Chinese Journal of Computers, 2021, 44(5): 921-936.
孙红帅, 王霞, 柳萱, 等. 频域注意力机制下的癫痫脑电信号分类 [J]. 西安交通大学学报, 2021, 55(2): 129-135.SUN Hongshuai, WANG Xia, LIU Xuan, et al. A classification method of epileptic electroencephalograms underfrequency-domain attention mechanism [J]. Journal of Xi'an Jiaotong University, 2021, 55(2): 129-135.
莫仁鹏, 李天梅, 司小胜, 等. 采用残差网络与卷积注意力机制的设备剩余使用寿命预测方法 [J]. 西安交通大学学报, 2022, 56(4): 194-202.
MO Renpeng, LI Tianmei, SI Xiaosheng, et al.Remaining useful life prediction for equipment using residual network and convolutional attention mechanism [J]. Journal of Xi'an Jiaotong University, 2022, 56(4): 194-202.
张连超, 乔瑞萍, 党祺玮, 等. 具有全局特征的空间注意力机制 [J]. 西安交通大学学报, 2020, 54(11): 129-138.
ZHANG Lianchao, QIAO Ruiping, DANG Qiwei, et al.Spatial attention mechanism with global characteristics [J]. Journal of Xi'an Jiaotong University, 2020, 54(11): 129-138.
WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module [C]//Computer Vision-ECCV 2018. Cham, Germany: Springer International Publishing, 2018: 3-19.
HOWARD A, SANDLER M, CHEN Bo, et al. Searching for MobileNetV3 [C]//2019 IEEE/CVF International Conference on Computer Vision. Piscataway, NJ, USA: IEEE, 2019: 1314-1324.
XIE Saining, GIRSHICK R, DOLLÁR P, et al. Aggregated residual transformations for deep neural networks [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2017: 5987-5995.
SANDLER M, HOWARD A, ZHU Menglong, et al. MobileNetV2: inverted residuals and linear bottlenecks [C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ, USA: IEEE, 2018: 4510-4520.
HU Jie, SHEN Li, ALBANIE S, et al. Squeeze-and-excitation networks [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42(8): 2011-2023.
SALAMON J, JACOBY C, BELLO J P. A dataset and taxonomy for urban sound research [C]//Proceedings of the 22nd ACM International Conference on Multimedia. New York, NY, USA: ACM, 2014: 1041-1044.
0
浏览量
5
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621