西安科技大学计算机科学与技术学院,西安,710054
: 2024-06-22。作者简介: 于振华(1977—),男,教授,博士生导师。基金项目: 国家自然科学基金资助项目(62273272)。
网络首发:2024-12-10,
纸质出版:2024
移动端阅览
于振华, 胡旭飞, 叶鸥. 类别条件生成对抗网络的语音对抗样本生成方法[J]. 西安交通大学学报, 2024,58(12):153-164.
YU Zhenhua, HU Xufei, YE Ou. Speech Adversarial Sample Generation Method Based on Class-Conditional Generative Adversarial Networks[J]. 2024, 58(12): 153-164.
于振华, 胡旭飞, 叶鸥. 类别条件生成对抗网络的语音对抗样本生成方法[J]. 西安交通大学学报, 2024,58(12):153-164. DOI: 10.7652/xjtuxb202412015.
YU Zhenhua, HU Xufei, YE Ou. Speech Adversarial Sample Generation Method Based on Class-Conditional Generative Adversarial Networks[J]. 2024, 58(12): 153-164. DOI: 10.7652/xjtuxb202412015.
针对现有面向自动语音识别系统的对抗攻击方法难以捕捉不同语音尺度之间的相关性、导致攻击成功率低的问题
提出一种类别条件生成对抗网络的语音对抗样本生成方法。通过目标标签映射模块
将攻击目标标签转化为独热向量
作为条件输入到构建的类别条件生成对抗网络中
以此控制语音样本类别的生成。该类别条件生成对抗网络中的生成器
采用设计的NReSidual U-block 网络模块与U-Net相融合
可以更好地学习不同时间尺度的语音特征
以及提升语音特征的表示能力
从而可以针对特定语音类别生成对抗样本; 判别器采用卷积块和全连接层相结合的网络结构
将错误损失通过梯度反向传播至生成器
能有效保留语音信号的时序信息
并解决数据分布不稳定问题。在通用的谷歌命令数据集和音乐流派数据集上进行实验
结果表明
所提语音对抗样本生成方法的攻击成功率与主流方法相比
分别提高了3.47%、5.1%
平均信噪比提升了3.2、1.49 dB
该方法具有较好的攻击效果和语音质量。
To address the challenge posed by existing adversarial attack methods for automatic speech recognition systems
which struggle to capture the correlation between different speech scales
resulting in a low attack success rate
a novel speech adversarial sample generation method based on class-conditional adversarial networks is introduced. Through the target label mapping module
the target label is transformed into a one-hot vector
serving as a conditional input to the constructed class-conditional generative adversarial network to control the generation of speech sample categories. The generator in this network combines the designed NReSidual U-block network module with U-Net to better learn speech features at different time scales and enhance the representation capability of speech features
thereby enabling the generation of adversarial samples tailored to specific speech categories. The discriminator adopts a network structure combining convolutional blocks and fully connected layers
enabling the propagation of error loss back to the generator through gradient backpropagation to effectively retain the temporal information of speech signals and address data distribution instability issues. Experiments conducted on the Google command dataset and music genre dataset demonstrate that the proposed speech adversarial sample generation method increases the attack success rate by 3.47% and 5.1%
respectively
compared to mainstream methods. Additionally
the average signal-to-noise ratio is improved by 3.2 and 1.49 dB. This method exhibits good attack effectiveness and speech quality.
LIU A H, HSU W N, AULI M, et al. Towards end-to-end unsupervised speech recognition [C]//2022 IEEE Spoken Language Technology Workshop. Piscataway, NJ, USA: IEEE, 2023: 221-228.
ZHAO Xia, WANG Limin, ZHANG Yufei, et al. A review of convolutional neural networks in computer vision [J]. Artificial Intelligence Review, 2024, 57(4): 99.
LÜ Qing, APIDIANAKI M, CALLISON-BURCH C. Towards faithful model explanation in NLP: a survey [J]. Computational Linguistics, 2024, 50(2): 657-723.
SUN Zheng, ZHAO Jinxiao, GUO Feng, et al. CommanderUAP: a practical and transferable universal adversarial attacks on speech recognition models [J]. Cybersecurity, 2024, 7(1): 38.
于振华, 殷正, 叶鸥, 等. 融合风格迁移的对抗样本生成方法 [J]. 西安交通大学学报, 2024, 58(7): 191-202.
YU Zhenhua, YIN Zheng, YE Ou, et al. Adversarial example generation method based on style transfer [J]. Journal of Xi'an Jiaotong University, 2024, 58(7): 191-202.
YUAN Xiaoyong, HE Pan, ZHU Qile, et al. Adversarial examples: attacks and defenses for deep learning [J]. IEEE Transactions on Neural Networks and Learning Systems, 2019, 30(9): 2805-2824.
ESMAEILPOUR M, CARDINAL P, KOERICH A L. Cyclic defense GAN against speech adversarial attacks [J]. IEEE Signal Processing Letters, 2021, 28: 1769-1773.
邹军华, 段晔鑫, 潘雨, 等. PIDI-FGSM:一种对抗样本生成的梯度处理新方法 [J]. 陆军工程大学学报, 2022, 1(5): 13-22.
ZOU Junhua, DUAN Yexin, PAN Yu, et al. Generating adversarial examples with PID iterative fast gradient sign method [J]. Journal of Army Engineering University of PLA, 2022, 1(5): 13-22.
李坤, 郭威, 张帆, 等. 基于遗传算法的恶意软件对抗样本生成方法 [J]. 计算机科学, 2023, 50(7): 325-331.
LI Kun, GUO Wei, ZHANG Fan, et al. Adversarial malware generation method based on genetic algorithm [J]. Computer Science, 2023, 50(7): 325-331.
KWON H, JEONG J. AdvU-net: generating adversarial example based on medical image and targeting U-net model [J]. Journal of Sensors, 2022, 2022(1): 4390413.
VAIDYA T, ZHANG Yuankai, SHERR M, et al. Cocaine noodles: exploiting the gap between human and machine speech recognition [C]//Proceedings of the 9th USENIX Conference on Offensive Technologies. New York, USA: USENIX Association, 2015: 16.
CISSE M, ADI Y, NEVEROVA N, et al. Houdini: fooling deep structured visual and speech recognition models with adversarial examples [C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc., 2017: 6980-6990.
AMODEI D, ANANTHANARAYANAN S, ANUBHAI R, et al. Deep speech 2: end-to-end speech recognition in English and Mandarin [C]//Proceedings of the 33rd International Conference on International Conference on Machine Learning. Chia Laguna Resort, Sardinia, Italy: PMLR, 2016: 173-182.
ESMAEILPOUR M, CARDINAL P, KOERICH A L.Towards robust speech-to-text adversarial attack [C]// IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2022: 2869-2873.
GUPTA H, WADHWA D S. Speech feature extraction and recognition using genetic algorithm [J]. International Journal of Emerging Technology and Advanced Engineering, 2014, 4(1): 363-369.
TAORI R, KAMSETTY A, CHU B, et al. Targeted adversarial examples for black box audio systems [C]//2019 IEEE Security and Privacy Workshops. Piscataway, NJ, USA: IEEE, 2019: 15-20.
DU Tianyu, JI Shouling, LI Jinfeng, et al. SirenAttack: generating adversarial audio for end-to-end acoustic systems [C]//Proceedings of the 15th ACM Asia Conference on Computer and Communications Security. New York, USA: Association for Computing Machinery, 2020: 357-369.
ZAGORUYKO S, KOMODAKIS N. Wide residual networks [C]//Proceedings of the British Machine Vision Conference. Dundee, UK: BMVA, 2016: 1-12.
CARLINI N, WAGNER D. Audio adversarial examples: targeted attacks on speech-to-text [C]//2018 IEEE Security and Privacy Workshops. Piscataway, NJ, USA: IEEE, 2018: 1-7.
NASSIF A B, SHAHIN I, ATTILI I, et al. Speech recognition using deep neural networks: a systematic review [J]. IEEE Access, 2019, 7: 19143-19165.
NEEKHARA P, HUSSAIN S, PANDEY P, et al. Universal adversarial perturbations for speech recognition systems [C]//Proceedings Interspeech 2019. Graz, Austria: ISCA, 2019: 481-485.
YUAN Xuejing, CHEN Yuxuan, ZHAO Yue, et al. Commandersong: a systematic approach for practical adversarial voice recognition [C]//Proceedings of the 27th USENIX Conference on Security Symposium. New York, USA: USENIX Association, 2018: 49-64.
XIE Yi, SHI Cong, LI Zhuohang, et al. Real-time, universal, and robust adversarial attacks against speaker recognition systems [C]// IEEE International Conference on Acoustics, Speech and Signal Processing. Piscataway, NJ, USA: IEEE, 2020: 1738-1742.
KONG J, KIM J, BAE J. HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis [C]//Proceedings of the 34th International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc., 2020: 17022-17033.
GAUTHIER J. Conditional generative adversarial nets for convolutional face generation [EB/OL]. [2024-05-06]. http://www.foldl.me/uploads/2015/conditional-gans-face-generation/paper.pdf.
MEHRA S, RANGA V, AGARWAL R. Improving speech command recognition through decision-level fusion of deep filtered speech cues [J]. Signal Image and Video Processing, 2024, 18(2): 1365-1373.
TZANETAKIS G, COOK P. Musical genre classification of audio signals [J]. IEEE Transactions on Speech and Audio Processing, 2002, 10(5): 293-302.
WANG Donghua, DONG Li, WANG Rangding, et al. Targeted speech adversarial example generation with generative adversarial network [J]. IEEE Access, 2020, 8: 124503-124513.
ZHANG Wenbin, LIAO Jun, ZHANG Yi, et al. CMGAN: a generative adversarial network embedded with causal matrix [J]. Applied Intelligence, 2022, 52(14): 16233-16245.[30] LIU Xin, ZHANG Weiwei, ZHENG Zhaohui, et al. FGP-GAN: fine-grained perception integrated generative adversarial network for expressive mandarin singing voice synthesis [J/OL]. IEEE Transactions on Consumer Electronics, 2024.(2024-06-11)[2024-05-06]. https://doi.org/10.1109/TCE.2024.3412053.
0
浏览量
25
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621