西北大学 计算机学院,710127,西安
收稿:2026-06-12,
修回:2026-08-22,
录用:2026-08-23,
移动端阅览
许祖荫, 孟宪佳, 解家豪. 面向智能驾驶的广义零样本驾驶行为识别方法[J]. 西安交通大学学报,2026.
XU Zuyin, MENG Xianjia, XIE Jiahao. A Generalized Zero-Shot Driving Behavior Recognition Method for Intelligent Driving[J]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY,2026.
针对智能驾驶行为识别依赖预定义类别、不可见行为缺少真实训练样本以及复杂生成模型难以边缘部署的问题,面向广义零样本学习任务,提出结合语义条件生成与潜在谱校准的驾驶行为识别方法(Diff-HAR)。首先,采用卷积时序编码器获得统一潜在表示,以类别语义为条件通过扩散模型生成不可见类特征;其次,利用主子空间保持的潜在谱校准和统计扰动调整生成特征结构与分布;最后,将其与真实可见类特征联合用于轻量分类器训练,并通过偏置校准实现统一识别,最终形成云端离线生成训练—边缘在线轻量推理的解耦架构。实验结果表明,在DriverActivity数据集中,Diff-HAR方法的不可见类准确率和调和平均值分别为39.0%和44.1%,较条件变分自编码器(CVAE)分别高10.1%和7.7%,与带梯度惩罚的Wasserstein生成对抗网络(WGAN-GP)的对应指标39.6%和43.9%总体相当;在PAMAP2数据集上,两项指标分别达到94.7%和72.3%,较FM-IoT分别高41.0%和10.2%;ESP32-S3轻量分类头回放测试中,模型大小为0.24 MB,平均推理延迟为9.379 ms。研究结果表明,Diff-HAR方法可在不可见类无真实训练样本的条件下兼顾可见类与不可见类识别,并验证了云端生成训练—边缘轻量分类在资源受限智能驾驶节点上的可行性。
To address the dependence of intelligent-driving behavior recognition on predefined classes
the lack of real training samples for unseen behaviors
and the difficulty of deploying complex generative models at the edge
a driving behavior recognition method combining semantic-conditioned generation and latent spectral calibration (Diff-HAR) is proposed for generalized zero-shot learning. First
a convolutional temporal encoder is used to obtain a unified latent representation
and unseen-class features are generated by a diffusion model conditioned on class semantics. Second
principal-subspace-preserving latent spectral calibration and statistical perturbation are used to adjust the structure and distribution of the generated features. Finally
the generated features are combined with real seen-class features to train a lightweight classifier
and bias calibration enables unified recognition
forming a decoupled architecture of cloud-based offline generation and training and edge-based online lightweight inference. Experimental results show that
on the DriverActivity dataset
Diff-HAR achieves an unseen-class accuracy of 39.0% and a harmonic mean of 44.1%
exceeding the conditional variational autoencoder (CVAE) by 10.1 and 7.7 percentage points
respectively
and is overall comparable to the Wasserstein generative adversarial network with gradient penalty (WGAN-GP)
whose corresponding metrics are 39.6% and 43.9%. On the PAMAP2 dataset
the two metrics reach 94.7% and 72.3%
respectively
outperforming FM-IoT by 41.0 and 10.2 percentage points. In the ESP32-S3 replay test of the lightweight classification head
the model size is 0.24 MB and the average inference latency is 9.379 ms. The results show that Diff-HAR balances seen- and unseen-class recognition without real training samples for unseen classes and demonstrate the feasibility of cloud-based generation and training with lightweight edge classification on resource-constrained intelligent-driving nodes.
黄昭彦 , 杨烁 , 吴建华 , 等 . 基于信息融合的智能网联汽车安全交互决策 [J ] . 自动化学报 , 2025 , 51 ( 9 ): 1883 - 1898 .
Huang Zhaoyan , Yang Shuo , Wu Jianhua , et al . Safety interactive decision-making for intelligent connected vehicles based on information fusion [J ] . Acta Automatica Sinica , 2025 , 51 ( 9 ): 1883 - 1898 .
张新钰 , 卢毅果 , 高鑫 , 等 . 面向智能网联汽车的车路协同感知技术及发展趋势 [J ] . 自动化学报 , 2025 , 51 ( 2 ): 233 - 248 .
Zhang Xinyu , Lu Yiguo , Gao Xin , et al . Vehicle-road collaborative perception technology and development trend for intelligent connected vehicles [J ] . Acta Automatica Sinica , 2025 , 51 ( 2 ): 233 - 248 .
丁飞 , 张楠 , 李升波 , 等 . 智能网联车路云协同系统架构与关键技术研究综述 [J ] . 自动化学报 , 2022 , 48 ( 12 ): 2863 - 2885 .
Ding Fei , Zhang Nan , Li Shengbo , et al . A survey of architecture and key technologies of intelligent connected vehicle-road-cloud cooperation system [J ] . Acta Automatica Sinica , 2022 , 48 ( 12 ): 2863 - 2885 .
陈妍妍 , 田大新 , 林椿眄 , 等 . 端到端自动驾驶系统研究综述 [J ] . 中国图象图形学报 , 2024 , 29 ( 11 ): 3216 - 3237 .
Chen Yanyan , Tian Daxin , Lin Chunmian , et al . Survey of end-to-end autonomous driving systems [J ] . Journal of Image and Graphics , 2024 , 29 ( 11 ): 3216 - 3237 .
董连飞 , 马志雄 , 朱西产 . 基于车载毫米波雷达动态手势识别网络 [J ] . 北京理工大学学报 , 2023 , 43 ( 5 ): 493 - 498 .
Dong Lianfei , Ma Zhixiong , Zhu Xichan . Dynamic gesture recognition network based on vehicular millimeter wave radar [J ] . Transactions of Beijing Institute of Technology , 2023 , 43 ( 5 ): 493 - 498 .
朱煜 , 赵江坤 , 王逸宁 , 等 . 基于深度学习的人体行为识别算法综述 [J ] . 自动化学报 , 2016 , 42 ( 6 ): 848 - 857 .
Zhu Yu , Zhao Jiangkun , Wang Yining , et al . A review of human action recognition based on deep learning [J ] . Acta Automatica Sinica , 2016 , 42 ( 6 ): 848 - 857 .
赵荣峰 , 卢宝莉 , 唐小江 , 等 . 面向智能座舱的多源混合模态数据集及层次化融合分类方法 [J ] . 智能系统学报 , 2026 , 21 ( 1 ): 83 - 94 .
Zhao Rongfeng , Lu Baoli , Tang Xiaojiang , et al . Multi-source hybrid-modality dataset and hierarchical fusion classification method for intelligent cockpits [J ] . CAAI Transactions on Intelligent Systems , 2026 , 21 ( 1 ): 83 - 94 .
胡宏宇 , 黎烨宸 , 张争光 , 等 . 基于多尺度骨架图和局部视觉上下文融合的驾驶员行为识别方法 [J ] . 汽车工程 , 2024 , 46 ( 1 ): 1 - 8 .
Hu Hongyu , Li Yechen , Zhang Zhengguang , et al . Driver behavior recognition based on multi-scale skeleton graph and local visual context method [J ] . Automotive Engineering , 2024 , 46 ( 1 ): 1 - 8 .
李少凡 , 高尚兵 , 张莹莹 . 用于驾驶员分心行为识别的姿态引导实例感知学习 [J ] . 中国图象图形学报 , 2023 , 28 ( 11 ): 3550 - 3561 .
Li Shaofan , Gao Shangbing , Zhang Yingying . Pose-guided instance-aware learning for driver distraction recognition [J ] . Journal of Image and Graphics , 2023 , 28 ( 11 ): 3550 - 3561 .
吕露露 , 黄毅 , 高君宇 , 等 . 多模态零样本人体动作识别 [J ] . 中国图象图形学报 , 2021 , 26 ( 7 ): 1658 - 1667 .
Lyu Lulu , Huang Yi , Gao Junyu , et al . Multimodal-based zero-shot human action recognition [J ] . Journal of Image and Graphics , 2021 , 26 ( 7 ): 1658 - 1667 .
Akata Z , Perronnin F , Harchaoui Z , et al . Label-embedding for image classification [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2016 , 38 ( 7 ): 1425 - 1438 .
Xian Y , Lorenz T , Schiele B , et al . Feature generating networks for zero-shot learning [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 2018 : 5542 - 5551 .
Chen S , Wang W , Xia B , et al . FREE: Feature refinement for generalized zero-shot learning [C ] // Proceedings of the IEEE/CVF International Conference on Computer Vision . 2021 : 122 - 131 .
Wang W , Li Q . Generalized zero-shot activity recognition with embedding-based method [J ] . ACM Transactions on Sensor Networks , 2023 , 19 ( 3 ): 1 - 25 .
Reimers N , Gurevych I . Sentence-BERT: Sentence embeddings using Siamese BERT-networks [C ] // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing . 2019 : 3982 - 3992 .
Devlin J , Chang M W , Lee K , et al . BERT: Pre-training of deep bidirectional transformers for language understanding [C ] // Proceedings of NAACL-HLT . 2019 : 4171 - 4186 .
Xue D , Fan X , Chen T , et al . Leveraging foundation models for zero-shot IoT sensing [C ] // Proceedings of the 27th European Conference on Artificial Intelligence . Amsterdam : IOS Press , 2024 : 3805 - 3812 .
Ji S , Zheng X , Wu C . HARGPT: Are LLMs zero-shot human activity recognizers? [C ] // 2024 IEEE International Workshop on Foundation Models for Cyber-Physical Systems & Internet of Things (FMSys) . 2024 : 38 - 43 .
Chowdhury R R , Kapila R , Panse A , et al . ZeroHAR: Sensor context augments zero-shot wearable action recognition [C ] // Proceedings of the AAAI Conference on Artificial Intelligence . 2025 , 39 ( 15 ): 16046 - 16054 .
Su J , Ge F , Wen Z , et al . IMUZero: Zero-shot human activity recognition by language-based cross modality fusion [J ] . Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , 2025 , 9 ( 4 ): 1 - 28 .
Ho J , Jain A , Abbeel P . Denoising diffusion probabilistic models [C ] // Advances in Neural Information Processing Systems . 2020 , 33 : 6840 - 6851 .
Oppel H , Munz M . A diffusion model for inertial based time series generation on scarce data availability to improve human activity recognition [J ] . Scientific Reports , 2025 , 15 : 16841 .
Ho J , Salimans T . Classifier-free diffusion guidance [EB/OL ] . arXiv : 2207 . 12598 , 2022 [ 2026-06-24 ] .
Reiss A , Stricker D . Introducing a new benchmarked dataset for activity monitoring [C ] // Proceedings of the 16th International Symposium on Wearable Computers . 2012 : 108 - 109 .
Zhang M , Sawchuk A A . USC-HAD: A daily activity dataset for ubiquitous activity recognition using wearable sensors [C ] // Proceedings of the 2012 ACM Conference on Ubiquitous Computing . 2012 : 1036 - 1043 .
Yang J , Huang H , Zhou Y , et al . MM-Fi: Multi-modal non-intrusive 4D human dataset for versatile wireless sensing [C ] // Advances in Neural Information Processing Systems . 2023 , 36 .
Li G H , Chiang H C , Li Y C , et al . A driver activity dataset with multiple RGB-D cameras and mmWave radars [C ] // Proceedings of the 15th ACM Multimedia Systems Conference . 2024 : 360 - 366 .
0
浏览量
1
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621