To realize high-speed motion of the multi-axis motion system and solve the problem of dynamic overload for motor
a time-optimal trajectory optimization scheme with reinforcement learning is proposed. This scheme chooses an improved state-action-reward-state-action(SARSA)algorithm and iterative interaction to find the time-optimal trajectory. The agent interacts with the reinforcement learning environment established by the kinematics model via the improved SARSA algorithm to find the initial trajectory satisfying the kinematics constraints. Then the agent interacts with the real environment by iterative interaction to introduce the motor dynamic load constraints into the reinforcement learning environment and modify the strategy trajectory so as to obtain the time-optimal trajectory satisfying motor dynamic load constraints. The time-optimal trajectory is verified on a self-built two-axis motion system. The results show that the speed and acceleration curves of the trajectory optimized by the improved SARSA algorithm are within the constraint ranges and the actual measured torque curve of the trajectory after 10 iterations is also within the dynamic load constraint ranges of motor
thus the time-optimal trajectory satisfying both kinematic constraints and dynamic load constraints of motor is obtained by this way.
关键词
Keywords
references
KAHN M E, ROTH B. The near-minimum-time control of open-loop articulated kinematic chains [J]. Journal of Dynamic Systems, Measurement, and Control, 1971, 93(3): 164-172.
PFEIFFER F, JOHANNI R. A concept for manipulator trajectory planning [J]. IEEE Journal on Robotics and Automation, 1987, 3(2): 115-123.
PHAM Q C. A general, fast, and robust implementation of the time-optimal path parameterization algorithm [J]. IEEE Transactions on Robotics, 2014, 30(6): 1533-1540.
SHILLER Z, LU H H. Computation of path constrained time-optimal motions with dynamic singularities [J]. Journal of Dynamic Systems, Measurement, and Control, 1992, 114(1): 34-40.
STEINHAUSER A, SWEVERS J. An efficient iterative learning approach to time-optimal path tracking for industrial robots [J]. IEEE Transactions on Industrial Informatics, 2018, 14(11): 5200-5207.
WALTZ M, FU K. A heuristic approach to reinforcement learning control systems [J]. IEEE Transactions on Automatic Control, 1965, 10(4): 390-398.
LOW E S, ONG P, CHEAH K C. Solving the optimal path planning of a mobile robot using improved Q-learning [J]. Robotics and Autonomous Systems, 2019, 115: 143-161.
KAREEM J M A, AL-ROUSAN M, QUADAN L. Reinforcement based mobile robot navigation in dynamic environment [J]. Robotics and Computer-Integrated Manufacturing, 2011, 27(1): 135-149.
KONAR A, GOSWAMI C I, SINGH S J, et al. A deterministic improved Q-learning for path planning of a mobile robot [J]. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2013, 43(5): 1141-1153.
WONG C C, LIU C C, XIAO S R, et al. Q-learning of straightforward gait pattern for humanoid robot based on automatic training platform [J]. Electronics, 2019, 8(6): 615.
ERDEN M S, LEBLEBICIOGLU K. Free gait generation with reinforcement learning for a six-legged robot [J]. Robotics and Autonomous Systems, 2008, 56(3): 199-212.
DUGULEANA M, BARBUCEANU F G, TEIRELBAR A, et al. Obstacle avoidance of redundant manipulators using neural networks based reinforcement learning [J]. Robotics and Computer-Integrated Manufacturing, 2012, 28(2): 132-146.
SANGIOVANNI B, RENDINIELLO A, INCREMONA G P, et al. Deep reinforcement learning for collision avoidance of robotic manipulators [C]∥Proceedings of the 2018 European Control Conference(ECC). Piscataway, NJ, USA: IEEE, 2018: 2063-2068.
INOUE T, DE MAGISTRIS G, MUNAWAR A, et al. Deep reinforcement learning for high precision assembly tasks [C]∥Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). Piscataway, NJ, USA: IEEE, 2017: 819-825.
REN Tianyu, DONG Yunfei, WU Dan, et al. Learning-based variable compliance control for robotic assembly [J]. Journal of Mechanisms and Robotics, 2018, 10(6): 061008.
SHIN K, MCKAY N. A dynamic programming approach to trajectory planning of robotic manipulators [J]. IEEE Transactions on Automatic Control, 1986, 31(6): 491-500.
XIAO Yongqiang, DU Zhijiang, DONG Wei. Smooth and near time-optimal trajectory planning of industrial robots for online applications [J]. Industrial Robot: An International Journal, 2012, 39(2): 169-177.
BOBROW J E, DUBOWSKY S, GIBSON J S. Time-optimal control of robotic manipulators along specified paths [J]. The International Journal of Robotics Research, 1985, 4(3): 3-17.
RUMMERY G A, NIRANJAN M. On-line Q-learning using connectionist system [EB/OL]. [2020-10-01]. http: ∥citeseer.ist.psu.edu/viewdoc/download; jsessionid=11751075B3B7CC74CF68CCE958731710?doi=10.1.1.17.2539rep=rep1type=pdf.
CHENG M Y, TSAI M C, KUO J C. Real-time NURBS command generators for CNC servo controllers [J]. International Journal of Machine Tools and Manufacture, 2002, 42(7): 801-813.