1.兰州理工大学自动化与电气工程学院,730050,兰州
2.兰州理工大学微电子现代产业学院,730050,兰州
收稿:2026-04-22,
修回:2026-07-27,
录用:2026-07-28,
移动端阅览
平梦梦, 李策, 唐广旭, 等. 面向具身智能体连续控制任务的可解释线性因果策略模型[J/OL]. 西安交通大学学报, 2026.
PING Mengmeng, LI Ce, TANG Guangxu, et al. An Interpretable Linear Causal Policy Model for Continuous Control Tasks of Embodied Agents[J/OL]. JOURNAL OF XI’AN JIAOTONG UNIVERSITY, 2026.
为解决具身智能体在高维状态与连续动作空间中的策略可解释性问题,提出了基于强化学习的状态–动作结构因果建模框架,利用结构因果模型显式刻画状态与动作依赖,提升策略的可解释性。该框架在仿真环境中训练并冻结策略网络,然后基于经验回放,构建离线状态–动作数据集;在此基础上,采用线性因果结构学习方法,对状态–动作中的因果结构进行显式建模;为提升因果结构学习的稳定性与泛化能力,引入策略敏感度驱动的加权稀疏先验与分组约束正则项,并结合稳定性筛选机制,通过抑制噪声与偶然相关性对结构推断的干扰,提升所学因果关系的性能。在9个MuJoCo动力学仿真平台具身连续控制任务上的实验结果表明:所提方法在平均奖励性能上比软演员-评论家提高了1.58%,中位数提高了0.38%;方法能够学习到稳定的因果结构,在与主流方法的对比中,多数任务表现出具有竞争力的性能。该研究在不增加额外交互成本的前提下实现了有效的因果策略建模,可为连续控制任务中强化学习策略的状态–动作因果分析提供参考。
To address the problem of policy interpretability for embodied agents in high-dimensional state and continuous action spaces
this study proposes a state–action structural causal modeling framework for RL policies
which aims to improve decision interpretability by explicitly characterizing state–action dependencies with structural causal models. The framework first trains and freezes a policy network in simulation
then constructs an offline state–action dataset from an experience replay buffer. Based on this dataset
we adopt a linear causal structure learning method to explicitly model the causal structure between states and actions. To further improve the stability and generalization of causal structure learning
we introduce a policy-sensitivity-driven weighted sparsity prior with group-regularized constraints
together with a stability selection mechanism. These components suppress the interference of noise and spurious correlations in structure inference
thereby enhancing the robustness of the learned causal relations. Experimental results on nine MuJoCo dynamics simulation platform embodied continuous control tasks show that the proposed method improves the average reward performance over soft actor-critic by 1.58% and the median performance by 0.38%
while learning stable causal structures. In comparison with mainstream methods
the proposed method demonstrates competitive performance across most tasks
indicating that it can maintain strong control capability while providing explicit causal interpretability. Therefore
this study achieves effective causal policy modeling without introducing additional interaction costs
providing methodological support for state–action causal analysis of reinforcement learning policies in continuous control tasks.
SEWAK M . Deep Reinforcement Learning [M ] . Springer , 2019 .
HENDERSON P , ISLAM R , BACHMAN P , et al . Deep Reinforcement Learning That Matters [C ] . Proceedings of the AAAI Conference on Artificial Intelligence , 2018 , 32 ( 1 ) : 3207 - 3214 .
LE N , RATHOUR V S , YAMAZAKI K , et al . Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey [J ] . Artificial Intelligence Review , 2022 , 55 ( 4 ): 2733 – 819 .
刘光远 , 曹晶仪 , 杜婕 . 联合遗传算法和强化学习的虚拟网络功能映射与调度方法 [J ] . 西安交通大学学报 , 2024 , 58 ( 8 ): 175 – 84 .
LIU G , CAO J , DU J . Function Mapping and Scheduling Method of Virtual Network Combining Genetic Algorithm and Reinforcement Learning [J ] . Journal of Xi’an Jiaotong Univerisity , 2024 , 58 ( 8 ): 175 – 84 .
METZ Y , GEISZL A , BAUR R , et al . Reward Learning from Multiple Feedback Types [C ] proceedings of the International Conference on Learning Representations , 2025 : 1 – 51 .
DONG J , HSU H-L , PAJIC M , et al . Efficiently Robust in-Context Reinforcement Learning with Adversarial Generalization and Adaptation [C ] . Proceedings of the Advances in Neural Information Processing Systems Workshop : Reliable ML from Unreliable Data , 2025 : 1 – 37 .
CALLAGHAN A , MASON K , MANNION P . Moma-Ac: A Preference-Driven Actor-Critic Framework for Continuous Multi-Objective Multi-Agent Reinforcement Learning [J ] . Neurocomputing , 2025 : 132032 .
LAN G , HAN D-J , HASHEMI A , et al . Asynchronous Federated Reinforcement Learning with Policy Gradient Updates: Algorithm Design and Convergence Analysis [C ] . Proceedings of the International Conference on LearningRepresentations , 2024 : 1 - 31 .
WANG Z , LIU J , PAN L . Learning Intractable Multimodal Policies with Reparameterization and Diversity Regularization [J ] . arXiv preprint 2025: arXiv: 2511.01374 .
马宇 , 安豆 , 林熙祥 , 等 . 面向飞行器智能协同控制的分层双时延策略梯度强化学习方法 [J ] . 西安交通大学学报 , 2025 , 59 ( 9 ): 88 – 98 .
MA Y , AN D , LIN X , et al . Hierarchical TwinDelayed Policy Gradient Reinforcement Learning for Intelligent Cooperative Control of Aircraft [J ] . Journal of Xi’an Jiaotong University , 2025 , 59 ( 9 ): 88 - 98 .
DELFOSSE Q , SHINDO H , DHAMI D , et al . Interpretable and Explainable Logical Policies Via Neurally Guided Symbolic Abstraction [J ] . Advances in Neural Information Processing Systems , 2023 , 36 : 50838 – 58 .
唐蕾 , 牛园园 , 王瑞杰 , 等 . 强化学习的可解释方法分类研究 [J ] . 计算机应用研究 , 2024 , 41 ( 6 ): 1601 – 1609 .
TANG L , NIU Y , WANG R , et al . Classification study of interpretable methods for reinforcement learning [J ] . Application Research of Computers , 2024 , 41 ( 6 ): 1601 - 1609 .
GLANOIS C , WENG P , ZIMMER M , et al . A Survey on Interpretable Reinforcement Learning [J ] . Machine Learning , 2024 , 113 ( 8 ): 5847 – 90 .
MOTT A , ZORAN D , CHRZANOWSKI M , et al . Towards Interpretable Reinforcement Learning Using Attention Augmented Agents [C ] . Proceedings of the Advances in Neural Information Processing Systems , 2019 , 32 : 1 – 10 .
TODOROV E , EREZ T , TASSA Y . Mujoco: A Physics Engine for Model-Based Control [C ] Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012 : 5026 – 33 .
HAARNOJA T , ZHOU A , ABBEEL P , et al . Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor [C ] . Proceedings of the International conference on machine learning , 2018 : 1861 – 70 .
ADAMCZYK J , MAKARENKO V , TIOMKIN S , et al . Average-Reward Soft Actor-Critic [C ] . Proceedings of the Reinforcement Learning Conference , 2025 : 1 – 19 .
XU Y , WEI Y , JIANG K , et al . Action Decoupled Sac Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces [J ] . Neurocomputing , 2023 , 537 : 141 – 51 .
WU Q , WANG Y , ZHAN S S , et al . Directly Forecasting Belief for Reinforcement Learning with Delays [C ] . Proceedings of the International Conference on Machine Learning , 2025 .
CHENG G , DONG L , CAI W , et al . Multi-Task Reinforcement Learning with Attention-Based Mixture of Experts [J ] . IEEE Robotics and Automation Letters , 2023 , 8 ( 6 ): 3812 – 3819 .
LIU S , ZHU M . Utility: Utilizing Explainable Reinforcement Learning to Improve Reinforcement Learning [C ] . Proceedings of the International Conference on Learning Representations , 2025 : 1 – 30 .
PROSPERI M , GUO Y , SPERRIN M , et al . Causal Inference and Counterfactual Prediction in Machine Learning for Actionable Healthcare [J ] . Nature Machine Intelligence , 2020 , 2 ( 7 ): 369 – 75 .
NAKANISHI T . Causalaime: Leveraging Peter-Clark Algorithms and Inverse Modeling for Unified Global Feature Explanation in Healthcare [C ] Proceedings of the World Conference on Explainable Artificial Intelligence , 2025 : 332 – 56 .
LEE K , RIBEIRO B , KOCAOGLU M . Constraint-Based Causal Discovery from a Collection of Conditioning Sets [C ] . Proceedings of the Conference on Uncertainty in Artificial Intelligence , 2025 : 1 – 31 .
CHEN Z , TIAN Z , ZHU J , et al . C-Cam: Causal Cam for Weakly Supervised Semantic Segmentation on Medical Image [C ] . Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022 : 11676 – 85 .
SUN Z , HAN Q , YANG H , et al . Invariant Deep Uplift Modeling for Incentive Assignment in Online Marketing Via Probability of Necessity and Sufficiency [C ] . Proceedings of the International Conference on Machine Learning , 2025 : 1 – 19 .
SUN Z , HAN Q , YANG H , et al . Invariant Deep Uplift Modeling for Incentive Assignment in Online Marketing Via Probability of Necessity and Sufficiency [C ] . Proceedings of the International Conference on Machine Learning , 2025 .
SPIRTES P , GLYMOUR C . An Algorithm for Fast Recovery of Sparse Causal Graphs [J ] . Social Science Computer Review , 1991 , 9 ( 1 ): 62 – 72 .
YI H , HE Y , CHEN D , et al . The Robustness of Differentiable Causal Discovery in Misspecified Scenarios [C ] . Proceedings of the International Conference on Learning Representations , 2025 : 1 – 39 .
孙悦雯 , 柳文章 , 孙长银 . 基于因果建模的强化学习控制: 现状及展望 [J ] . 自动化学报 , 2023 , 49 ( 3 ): 661 – 77 .
SUN Y-W , LIU W-Z , SUN C-Y . Causalityin Reinforcement Learning Control: The State of the Art and Prospects . Acta Automatica Sinica , 2023 , 49 ( 3 ): 661 − 677 .
KARLSSON R K A , KRIJTHE J H . Falsification of Unconfoundedness by Testing Independence of Causal Mechanisms [C ] . Proceedings of the International Conference on Machine Learning , 2025 : 1 - 20 .
SUN X , SCHULTE O , LIU G , et al . Nts-Notears: Learning Nonparametric Dbns with Prior Knowledge [C ] . Proceedings of the International Conference on Artificial Intelligence and Statistics , 2023 : 1 - 23 .
SCHULMAN J , WOLSKI F , DHARIWAL P , et al . Proximal Policy Optimization Algorithms [J ] . arXiv preprint arXiv: 1707.06347 , 2017 .
FUJIMOTO S , HOOF H , MEGER D . Addressing Function Approximation Error in Actor-Critic Methods [C ] . Proceedings of the International Conference on Machine Learning. PMLR , 2018 : 1587 - 1596 .
NAZARET A , HONG J , AZIZI E , et al . Stable Differentiable Causal Discovery [J ] . arXiv preprint arXiv: 2311.10263 , 2023 .
XU Z , LI Y , LIU C , et al . Ordering-based Causal Discovery for Linear and Nonlinear Relations [C ] . Proceedings of the Advances in Neural Information Processing Systems , 2024 , 37 : 4315 - 4340 .
DUONG B , GUPTA S , NGUYEN T . Causal Discovery via Bayesian Optimization [J ] . arXiv preprint arXiv: 2501.14997 , 2025 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621