华南理工大学机械与汽车工程学院,广州,510641
网络首发:2021-07-10,
纸质出版:2021
移动端阅览
张铁, 廖才磊, 邹焱飚, 等. 采用强化学习的多轴运动系统时间最优轨迹优化[J]. 西安交通大学学报, 2021,55(7):33-40.
Time-Optimal Trajectory Optimization of Multi-Axis Motion System by Reinforcement Learning[J]. 2021, 55(7): 33-40.
张铁, 廖才磊, 邹焱飚, 等. 采用强化学习的多轴运动系统时间最优轨迹优化[J]. 西安交通大学学报, 2021,55(7):33-40. DOI: 10.7652/xjtuxb202107004.
Time-Optimal Trajectory Optimization of Multi-Axis Motion System by Reinforcement Learning[J]. 2021, 55(7): 33-40. DOI: 10.7652/xjtuxb202107004.
为实现多轴运动系统高速运动并解决电机动载荷过载的问题
提出了一种采用强化学习的时间最优轨迹优化方法。使用改进状态-动作-奖励-状态-动作(SARSA)算法和迭代交互法来寻找时间最优轨迹:通过改进SARSA算法与基于运动学模型建立的强化学习环境进行交互学习
找到满足运动学约束的初始策略轨迹; 通过迭代交互法与真实环境进行交互学习
从而将电机动态载荷约束引入到强化学习环境中并对策略轨迹进行修正; 最终得到满足电机动态载荷约束的时间最优轨迹。在自行搭建的两轴运动系统上进行验证
结果表明
改进SARSA算法优化得到的策略轨迹的速度和加速度曲线均在约束范围内
且经过10次迭代后的轨迹实际测量力矩曲线也在电机动载荷约束范围内
所提方法能够得到同时满足运动学约束和动力学约束的时间最优轨迹。
To realize high-speed motion of the multi-axis motion system and solve the problem of dynamic overload for motor
a time-optimal trajectory optimization scheme with reinforcement learning is proposed. This scheme chooses an improved state-action-reward-state-action(SARSA)algorithm and iterative interaction to find the time-optimal trajectory. The agent interacts with the reinforcement learning environment established by the kinematics model via the improved SARSA algorithm to find the initial trajectory satisfying the kinematics constraints. Then the agent interacts with the real environment by iterative interaction to introduce the motor dynamic load constraints into the reinforcement learning environment and modify the strategy trajectory so as to obtain the time-optimal trajectory satisfying motor dynamic load constraints. The time-optimal trajectory is verified on a self-built two-axis motion system. The results show that the speed and acceleration curves of the trajectory optimized by the improved SARSA algorithm are within the constraint ranges and the actual measured torque curve of the trajectory after 10 iterations is also within the dynamic load constraint ranges of motor
thus the time-optimal trajectory satisfying both kinematic constraints and dynamic load constraints of motor is obtained by this way.
KAHN M E, ROTH B. The near-minimum-time control of open-loop articulated kinematic chains [J]. Journal of Dynamic Systems, Measurement, and Control, 1971, 93(3): 164-172.
PFEIFFER F, JOHANNI R. A concept for manipulator trajectory planning [J]. IEEE Journal on Robotics and Automation, 1987, 3(2): 115-123.
PHAM Q C. A general, fast, and robust implementation of the time-optimal path parameterization algorithm [J]. IEEE Transactions on Robotics, 2014, 30(6): 1533-1540.
SHILLER Z, LU H H. Computation of path constrained time-optimal motions with dynamic singularities [J]. Journal of Dynamic Systems, Measurement, and Control, 1992, 114(1): 34-40.
STEINHAUSER A, SWEVERS J. An efficient iterative learning approach to time-optimal path tracking for industrial robots [J]. IEEE Transactions on Industrial Informatics, 2018, 14(11): 5200-5207.
WALTZ M, FU K. A heuristic approach to reinforcement learning control systems [J]. IEEE Transactions on Automatic Control, 1965, 10(4): 390-398.
LOW E S, ONG P, CHEAH K C. Solving the optimal path planning of a mobile robot using improved Q-learning [J]. Robotics and Autonomous Systems, 2019, 115: 143-161.
KAREEM J M A, AL-ROUSAN M, QUADAN L. Reinforcement based mobile robot navigation in dynamic environment [J]. Robotics and Computer-Integrated Manufacturing, 2011, 27(1): 135-149.
KONAR A, GOSWAMI C I, SINGH S J, et al. A deterministic improved Q-learning for path planning of a mobile robot [J]. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2013, 43(5): 1141-1153.
WONG C C, LIU C C, XIAO S R, et al. Q-learning of straightforward gait pattern for humanoid robot based on automatic training platform [J]. Electronics, 2019, 8(6): 615.
ERDEN M S, LEBLEBICIOGLU K. Free gait generation with reinforcement learning for a six-legged robot [J]. Robotics and Autonomous Systems, 2008, 56(3): 199-212.
DUGULEANA M, BARBUCEANU F G, TEIRELBAR A, et al. Obstacle avoidance of redundant manipulators using neural networks based reinforcement learning [J]. Robotics and Computer-Integrated Manufacturing, 2012, 28(2): 132-146.
SANGIOVANNI B, RENDINIELLO A, INCREMONA G P, et al. Deep reinforcement learning for collision avoidance of robotic manipulators [C]∥Proceedings of the 2018 European Control Conference(ECC). Piscataway, NJ, USA: IEEE, 2018: 2063-2068.
INOUE T, DE MAGISTRIS G, MUNAWAR A, et al. Deep reinforcement learning for high precision assembly tasks [C]∥Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). Piscataway, NJ, USA: IEEE, 2017: 819-825.
REN Tianyu, DONG Yunfei, WU Dan, et al. Learning-based variable compliance control for robotic assembly [J]. Journal of Mechanisms and Robotics, 2018, 10(6): 061008.
SHIN K, MCKAY N. A dynamic programming approach to trajectory planning of robotic manipulators [J]. IEEE Transactions on Automatic Control, 1986, 31(6): 491-500.
XIAO Yongqiang, DU Zhijiang, DONG Wei. Smooth and near time-optimal trajectory planning of industrial robots for online applications [J]. Industrial Robot: An International Journal, 2012, 39(2): 169-177.
BOBROW J E, DUBOWSKY S, GIBSON J S. Time-optimal control of robotic manipulators along specified paths [J]. The International Journal of Robotics Research, 1985, 4(3): 3-17.
RUMMERY G A, NIRANJAN M. On-line Q-learning using connectionist system [EB/OL]. [2020-10-01]. http: ∥citeseer.ist.psu.edu/viewdoc/download; jsessionid=11751075B3B7CC74CF68CCE958731710?doi=10.1.1.17.2539rep=rep1type=pdf.
CHENG M Y, TSAI M C, KUO J C. Real-time NURBS command generators for CNC servo controllers [J]. International Journal of Machine Tools and Manufacture, 2002, 42(7): 801-813.
张腾,张小栋,张英杰,陆竹风,朱文静,蒋永玉.引入深度强化学习思想的脑-机协作精密操控方法.2021,55(2):1-9.[doi:10.7652/xjtuxb202102001]
孙平,丁雨姗.全方向康复步行训练机器人具有死区补偿的反步有限时间控制.2020,54(7):1-8,74.[doi:10.7652/xjtuxb202007001]
张亮修,张铁柱,吴光强.考虑误差校正的智能车辆路径跟踪鲁棒预测控制.2020,54(3):20-27.[doi:10.7652/xjtuxb 202003003]
颛孙少帅,杨俊安,刘辉,黄科举.未知拓扑无线自组网络多节点干扰决策算法.2018,52(6):91-97.[doi:10.7652/xjtuxb201806015]
张铁,林康宇,邹焱飚,刘晓刚.用于机器人末端残余振动控制的控制误差优化输入整形器.2018,52(4):90-97.[doi:10.7652/xjtuxb201804014]
颛孙少帅,杨俊安,刘辉,黄科举.采用双层强化学习的干扰决策算法.2018,52(2):63-69.[doi:10.7652/xjtuxb201802 011]
0
浏览量
5
下载量
1
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621