1.吉林大学 通信工程学院,130022,吉林长春
2.吉林大学 汽车底盘集成与仿生全国重点实验室,130025,吉林长春
收稿:2026-06-30,
修回:2026-09-06,
录用:2026-09-08,
移动端阅览
胡云峰, 张茗客, 麻斌. 视觉语言模型的自动提示注入排版攻击方法[J]. 西安交通大学学报,2026.
HU Yunfeng, ZHANG Mingke, MA Bin. Automated Typographic Prompt Injection Attack on Vision-Language Models[J]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY,2026.
针对当前视觉语言模型多模态暴露出的安全脆弱性,本文提出了一套针对自动驾驶场景的自动提示注入排版攻击,构建了一条从驾驶场景感知、自动生成对抗语义、物理自适应渲染到跨模态决策劫持的端到端闭环链路。首先,利用视觉语言模型感知原始驾驶场景状态,通过大语言模型自动生成具有强语义翻转和诱导性的攻击文本指令;其次,结合目标检测模型,将攻击文本自适应地渲染并排版至画面中的合理物理载体上;最后,引入云端大语言模型作为自动化裁判,通过对比攻击前后的感知与规划报告,精准计算导致危险决策的攻击成功率。结果表明:在Qwen-VL、InternVL、BLIP和自动驾驶专精模型DriveMM四种模型上,本文所设计的自动提示注入排版攻击能够成功欺骗包括Qwen等各类先进视觉语言模型,攻击成功率分别为67.55%、61.51%、30.57%、10.38%。
In response to the security vulnerabilities exposed by current vision-language models in multimodal scenarios
this study proposes an Automated Typographic Prompt Injection Attack tailored for autonomous driving scenarios. The proposed method constructs an end-to-end closed-loop pipeline consisting of autonomous driving scene perception
automated adversarial semantic generation
physically adaptive rendering
and cross-modal decision hijacking. Specifically
a Vision-Language Model (VLM) is first employed to perceive the original driving scene state
after which a Large Language Model (LLM) automatically generates adversarial text instructions with strong semantic inversion and deceptive effects. Subsequently
combined with an object detection model
the generated attack texts are adaptively rendered and typeset onto appropriate physical carriers within the scene. Finally
a cloud-based Large Language Model is introduced as an automated referee to accurately evaluate the attack success rate of inducing hazardous decisions by comparing the perception and planning reports before and after the attack.Experimental results demonstrate that the proposed Automated Typographic Prompt Injection Attack successfully deceives various advanced Vision-Language Models
including Qwen-VL
InternVL
BLIP
and the autonomous driving specialized model DriveMM
achieving attack success rates of 67.55%
61.51%
30.57%
and 10.38%
respectively.
HU Y , YANG J , CHEN L , et al . Planning-oriented autonomous driving [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2023 : 17853 - 17862 .
杜少毅 , 孙远 , 刘宇颖 , 等 . 面向自动驾驶协同感知的高效传输方法综述 [J/OL ] . 西安交通大学学报 , 2026 .
DU Shaoyi , SUN Yuan , LIU Yuying , et al . A Review of Efficient Transmission for Cooperative Perception in Autonomous Driving [J/OL ] . JOURNAL OF XI’AN JIAOTONG UNIVERSITY , 2026 .
CUI C , MA Y , CAO X , et al . A survey on multimodal large language models for autonomous driving [C ] // Proceedings of the IEEE/CVF winter conference on applications of computer vision . 2024 : 958 - 979 .
LIU H , LI C , WU Q , et al . Visual instruction tuning [J ] . Advances in neural information processing systems , 2023 , 36 : 34892 - 34916 .
刘毅 , 秦贵和 , 赵睿 . 车载控制器局域网络安全协议 [J ] . 西安交通大学学报 , 2018 ( 5 ): 94 - 100 .
LIU Yi , QIN Guihe , ZHAO Rui . Security Protocol for On-Board Controller Area Network [J ] . JOURNAL OF XI'AN JIAOTONG UNIVERSITY , 2018 ( 5 ): 94 - 100 .
YIN H , ZHAO Z , YAN J , et al . Certified Robustness in Automated Driving Perception: A Review: H. Yin et al [J ] . Automotive Innovation , 2025 , 8 ( 4 ): 817 - 837 .
WANG X , JI Z , MA P , et al . Instructta: Instruction-tuned targeted attack for large vision-language models [J ] . arXiv preprint arXiv: 2312.01886 , 2023 .
LUO H , GU J , LIU F , et al . An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models [J ] . arXiv preprint arXiv: 2403.09766 , 2024 .
CUI X , APARCEDO A , JANG Y K , et al . On the robustness of large multimodal models against image adversarial attacks [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2024 : 24625 - 24634 .
王耀宁 , 王郅瑞 , 李云 . 基于特征对齐的大型视觉语言模型攻击迁移性研究 [J/OL ] . 计算机工程 , 1 - 10 [2026-06-28 ]
WANG Yaoning , WANG Zhirui , LI Yun . Enhancing Transferability of Adversarial Attacks on Large Vision-Language Models via Intermediate Layer Feature Alignment [J ] . Computer Engineering
GRESHAKE K , ABDELNABI S , MISHRA S , et al . Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection [C ] // Proceedings of the 16th ACM workshop on artificial intelligence and security . 2023 : 79 - 90 .
PAPERNOT N , MCDANIEL P , Goodfellow I . Transferability in machine learning: from phenomena to black-box attacks using adversarial samples [J ] . arXiv preprint arXiv: 1605.07277 , 2016 .
CHENG H , XIAO E , GU J , et al . Unveiling typographic deceptions: Insights of the typographic vulnerability in large vision-language models [C ] // European Conference on Computer Vision . Cham : Springer Nature Switzerland , 2024 : 179 - 196 .
LI J , LI D , XIONG C , et al . Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation [C ] // International conference on machine learning . PMLR , 2022 : 12888 - 12900 .
BAI J , BAI S , CHU Y , et al . Qwen technical report [J ] . arXiv preprint arXiv: 2309.16609 , 2023 .
CHEN Z , WU J , WANG W , et al . Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks [C ] // Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2024 : 24185 - 24198 .
GOH G , CAMMARATA N , VOSS C , et al . Multimodal neurons in artificial neural networks [J ] . Distill , 2021 , 6 ( 3 ): e30 .
RADFORD A , KIM J W , HALLACY C , et al . Learning transferable visual models from natural language supervision [C ] // International conference on machine learning . PmLR , 2021 : 8748 - 8763 .
QRAITEM M , TASNIM N , TETERWAK P , et al . Vision-llms can fool themselves with self-generated typographic attacks [J ] . arXiv preprint arXiv: 2402.00626 , 2024 .
CHUNG N , GAO S , VU T A , et al . Towards transferable attacks against vision-llms in autonomous driving with typography [J ] . arXiv preprint arXiv: 2405.14169 , 2024 .
OORD A , LI Y , VINYALS O . Representation learning with contrastive predictive coding [J ] . arXiv preprint arXiv: 1807.03748 , 2018 .
DOSOVITSKLY A , BEYER L , KOLESNIKOV A , et al . An image is worth 16x16 words: Transformers for image recognition at scale [J ] . arXiv preprint arXiv: 2010.11929 , 2020 .
BEHRENDT K , NOVAK L , BOTROS R . A deep learning approach to traffic lights: Detection, tracking, and classification [C ] // 2017 IEEE International Conference on Robotics and Automation (ICRA) . IEEE , 2017 : 1370 - 1377 .
Polley N , Pavlitska S , Boualili Y , et al . TLD-READY: Traffic Light Detection‐Relevance Estimation and Deployment Analysis [C ] // 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC) . IEEE , 2024 : 3800 - 3806 .
Pon A , Adrienko O , Harakeh A , et al . A hierarchical deep architecture and mini-batch selection method for joint traffic sign and light detection [C ] // 2018 15th Conference on Computer and Robot Vision (CRV) . IEEE , 2018 : 102 - 109 .
CAO Y , XING Y , ZHANG J , et al . Scenetap: Scene-coherent typographic adversarial planner against vision-language models in real-world environments [C ] // Proceedings of the Computer Vision and Pattern Recognition Conference . 2025 : 25050 - 25059
HUANG Z , FENG C , YAN F , et al . Drivemm: All-in-one large multimodal model for autonomous driving [J ] . arXiv preprint arXiv:2412.07689, 2024 , 2 ( 3 ): 8 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621