Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Reinforcement learning based constrained optimal control:An interpretable reward design

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Reinforcement learning based constrained optimal control:An interpretable reward design

Similar Papers
  • Research Article
  • Cite Count Icon 10
  • 10.1109/mcs.2014.2308694
Optimal Adaptive Control and Differential Games by Reinforcement Leanring Principles [Book review
  • Jun 1, 2014
  • IEEE Control Systems
  • Warren E Dixon

This book introduces and leverages concepts from reinforcement learning (RL). RL methods have been formalized by the computational intelligence community based on the conditioned reflex concept and serve as the bridge between adaptive and optimal control methods. In this book, the authors focus on RL methods to design adaptive controllers that learn approximate optimal solution methods. These optimal adaptive control approaches make use of neural networks (NNs) that serve as the actor-critic for discrete and continuous optimal control and differential game problems. The book contains 11 chapters, with most chapters representing the results from recent journal publications. For readers with an understanding of concepts such as Lyapunov-based analysis and adaptive and optimal control methods, this book presents a foundation for online optimal control solutions for continuous systems. The text uses RL to merge concepts from optimal and adaptive control, and the results are solutions to a class of challenging and highly relevant optimal control problems. The text is the first of its kind specialized for the control systems community. The mixed use of typical state feedback and sampled past data to adjust the adaptive update laws presents an exciting new domain rich for new contributions. As control practitioners seek to apply the approaches in this text to more complex, higher-dimensional systems, new challenges will likely arise such as finding appropriate basis functions for the learning laws; developing alternatives to or ensuring persistence of excitation, extending to interconnected (potentially distributed) networks; and applying to systems where the uncertainty enters the dynamics in unique ways. In summary, this book presents an exciting blend of ideas and methods developed in machine learning and adaptive control to provide new inroads into online approximate optimal control for complex and potentially uncertain systems.

  • Research Article
  • 10.1002/oca.2979
Data‐based learning control for optimization of nonlinear systems
  • Mar 9, 2023
  • Optimal Control Applications and Methods
  • Qinglai Wei + 3 more

With the development of science and technology, practical systems such as the power systems, traffic systems, robot manipulator systems, etc., have become more complex. Therefore, it is difficult to build practical systems by accurate models. Under the lack of accurate process models, using system data to improve system performance and learn optimal decisions becomes very important. Through the recent years, data-based learning control theories and technologies have widely been investigated, including adaptive dynamic programming, reinforcement learning, iterative learning control, and so on. Data-based methods require the system data instead of the accurate knowledge of system dynamics that can be considered as model-free learning control methods. The data-based methods are effective solutions for the optimal control of nonlinear systems, which motivate this special issue. This special issue aims to collect and present original research dealing with data-based learning and their applications for optimization and control problems. The first group of papers1-7 focuses on data-based control theory, approaches, and applications. A fuzzy model predictive control approach is proposed for stick-slip type piezoelectric actuator to realize the precise control of the end effector.1 A systematic online adaptive dynamic programming control framework is proposed for smart buildings control to ensure hard constraints to be satisfied.2 A multi-verse optimizer tuned PI-type active disturbance rejection generalized predictive control method is described for the motion control problems of ships.3 The sufficient optimality conditions for the optimal controls are established under some convexity assumptions.4 A receding-horizon reinforcement learning algorithm is proposed for near-optimal control of continuous-time systems under control constraints.5 In order to solve the interference compensation control problem of a class of nonlinear systems, a method based on memory data is introduced to suppress interference greatly.6 A new controller design method is proposed for the trajectory tracking problem of robots with imprecise dynamic properties and interference.7 The second group of papers8-12 considers iterative learning identification and iterative learning control. An iterative learning control approach is proposed for linear parabolic distributed parameter systems with multiple actuators and multiple sensors.8 The quantized data-based iterative learning tracking control problem is studied for nonlinear networked control systems with signals quantization and denial-of-service attacks.9 The output tracking problem is considered for a class of nonlinear parabolic distributed parameter systems with moving boundaries.10 A just-in-time learning based dual heuristic programming algorithm is proposed to optimize the control performance of autonomous wheeled mobile robots under faults or disturbances.11 A novel optimal constraint-following controller is proposed for uncertain mechanical systems.12 The third group of papers13-19 focuses on robustness on data-based optimal learning control. A novel Nash game-theoretical optimal adaptive robust control design approach is proposed to address the constraint-following control problem for the uncertain underactuated mechanical systems with fuzzy evidence theory.13 A partial model-free sliding mode control strategy is proposed for a class of disturbed systems.14 A new data-based adaptive dynamic programming algorithm is proposed to solve the optimal control policy for discrete-time systems with uncertainties.15 A method that applies event-triggered mechanism H ∞ $$ {\mathrm{H}}_{\infty } $$ control to continuous-time nonlinear systems with asymmetric constraints based on dual heuristic dynamic programming structure is proposed.16 A novel anti-disturbance inverse optimal controller design method is proposed for a class of high-dimensional chain structure systems with any disturbances, matched, or mismatched.17 A data-driven H ∞ $$ {\mathrm{H}}_{\infty } $$ controller design method is studied for continuous-time linear periodic systems.18 The problem of the post-stall pitching maneuver of an aircraft with lower deflection frequency of control actuator is studied by considering the unsteady aerodynamic disturbances.19 The fourth group of papers20-23 focuses on neural networks and deep neural networks learning methods for optimal control. An optimal tracking control problem for the injection flow front position arising in the filling process in the injection molding machine is considered, and an intelligent real-time optimal control method based on deep neural networks is developed for the online tracking of the flow front position to improve the efficient production process of the plastics.20 An efficient and systematic method is proposed for model-based predictive control synthesis.21 The decentralized control issues of nonlinear large-scale systems are investigated via critic-only adaptive dynamic programming learning methods.22 A singularity-free online neural network-based sliding mode control method is proposed to realize the fixed-wing perch maneuver.23 The fifth group of papers24-27 discusses data-based control for distributed control systems. A mission-driven control scheme, including a consensus-based near-optimal formation controller and a finite-time precise formation controller, is proposed aiming at different requirements of unmanned aerial vehicle swarm.24 The neural network adaptive formation control of a class of second-order nonlinear systems with unmodeled dynamics is investigated, where the control law merely depends on the relative bearings between neighboring agents.25 The neighbor Q-learning based consensus control algorithm is developed for discrete-time multiagent systems.26 The fault-tolerate containment control problem is considered for stochastic nonlinear multiagent systems in the presence of input saturation and sensor faults.27 The sixth group of papers28-30 considers applications of data-based learning methods to industrial processes. A stochastic gradient algorithm based on the minimum Shannon entropy is proposed to identify a type of Hammerstein system with random noise.28 A predictive control strategy based on Hammerstein–Wiener inverse model compensation is proposed aiming at the nonlinearity and large lag of the pH change in wet flue gas desulfurization process.29 An algorithm called the kernel entropy regression is proposed to enhance the interpretability between the fault and the key performance indicator.30 The seventh group of papers31-36 focuses on machine learning, data mining, and practical applications in automation. The performance of a Takagi–Sugeno fuzzy-model-based observer is enhanced by proposing a featured multi-instant united switch-type observer.31 The reinforcement learning theory with deep Q-network is applied for the mobile robot to achieve a collision-free path in an unknown dynamic environment.32 An energy-saving velocity planning algorithm is proposed for rail transit train with running and computation delays.33 A novel COVID-19 transmission model is established by introducing traditional susceptible–exposed–infected–removed disease transmission models into complex network.34 A novel collaborative diagnosis method is presented by combining variational modal decomposition and stochastic configuration network for incipient faults of rolling bearing.35 The linear dependence graph associated with a finite-dimensional vector space is studied.36 In summary, this special issue provides an opportunity to review the most recent developments in data-based learning control for optimization of nonlinear systems, by considering theory, algorithms, and applications.

  • Research Article
  • Cite Count Icon 93
  • 10.1109/tnnls.2022.3214681
Adaptive Optimal Tracking Control of an Underactuated Surface Vessel Using Actor-Critic Reinforcement Learning.
  • Jun 1, 2024
  • IEEE transactions on neural networks and learning systems
  • Lin Chen + 2 more

In this article, we present an adaptive reinforcement learning optimal tracking control (RLOTC) algorithm for an underactuated surface vessel subject to modeling uncertainties and time-varying external disturbances. By integrating backstepping technique with the optimized control design, we show that the desired optimal tracking performance of vessel control is guaranteed due to the fact that the virtual and actual control inputs are designed as optimized solutions of every subsystem. To enhance the robustness of vessel control systems, we employ neural network (NN) approximators to approximate uncertain vessel dynamics and present adaptive control technique to estimate the upper boundedness of external disturbances. Under the reinforcement learning framework, we construct actor-critic networks to solve the Hamilton-Jacobi-Bellman equations corresponding to subsystems of surface vessel to achieve the optimized control. The optimized control algorithm can synchronously train the adaptive parameters not only for actor-critic networks but also for NN approximators and adaptive control. By Lyapunov stability theorem, we show that the RLOTC algorithm can ensure the semiglobal uniform ultimate boundedness of the closed-loop systems. Compared with the existing reinforcement learning control results, the presented RLOTC algorithm can compensate for uncertain vessel dynamics and unknown disturbances, and obtain the optimized control performance by considering optimization in every backstepping design. Simulation studies on an underactuated surface vessel are given to illustrate the effectiveness of the RLOTC algorithm.

  • Research Article
  • Cite Count Icon 2
  • 10.1049/iet-cta.2016.0646
Guest Editorial
  • Aug 1, 2016
  • IET Control Theory & Applications
  • Jinliang Ding + 4 more

Guest Editorial

  • Research Article
  • Cite Count Icon 1
  • 10.1109/taes.2025.3588483
Prescribed Performance Optimal Backstepping Control for Hypersonic Vehicles Based on Reinforcement Learning With Disturbance Observer
  • Oct 1, 2025
  • IEEE Transactions on Aerospace and Electronic Systems
  • Haoyu Cheng + 4 more

This paper presents a prescribed performance optimal backstepping control method based on reinforcement learning (RL) to address the challenge of achieving optimal control for hypersonic vehicles (HSVs) while accounting for dynamic performance under disturbances. A nonlinear model of the hypersonic vehicle is developed, and the controller design is divided into altitude and velocity subsystems. Optimal control commands for each subsystem are derived using RL within the Actor-Critic framework. To enhance the anti-disturbance capabilities of RL, the total system uncertainties are estimated through an extended state observer (ESO), effectively balancing optimal control with disturbance rejection. The proposed lightweight RL method incorporates an adaptive update law for weight adjustment, eliminating the need for the trial-and-error process typical of conventional RL. Furthermore, a Time-varying Barrier Lyapunov Functions (TvBLFs) based on prescribed performance theory ensures the stability of the closed-loop system and ensures global state convergence within prescribed performance bounds. Simulation results confirm the effectiveness and superiority of the proposed method.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 8
  • 10.1002/acs.3115
Online optimal and adaptive integral tracking control for varying discrete‐time systems using reinforcement learning
  • Apr 16, 2020
  • International Journal of Adaptive Control and Signal Processing
  • Ibrahim Sanusi + 3 more

SummaryConventional closed‐form solution to the optimal control problem using optimal control theory is only available under the assumption that there are known system dynamics/models described as differential equations. Without such models, reinforcement learning (RL) as a candidate technique has been successfully applied to iteratively solve the optimal control problem for unknown or varying systems. For the optimal tracking control problem, existing RL techniques in the literature assume either the use of a predetermined feedforward input for the tracking control, restrictive assumptions on the reference model dynamics, or discounted tracking costs. Furthermore, by using discounted tracking costs, zero steady‐state error cannot be guaranteed by the existing RL methods. This article therefore presents an optimal online RL tracking control framework for discrete‐time (DT) systems, which does not impose any restrictive assumptions of the existing methods and equally guarantees zero steady‐state tracking error. This is achieved by augmenting the original system dynamics with the integral of the error between the reference inputs and the tracked outputs for use in the online RL framework. It is further shown that the resulting value function for the DT linear quadratic tracker using the augmented formulation with integral control is also quadratic. This enables the development of Bellman equations, which use only the system measurements to solve the corresponding DT algebraic Riccati equation and obtain the optimal tracking control inputs online. Two RL strategies are thereafter proposed based on both the value function approximation and the Q‐learning along with bounds on excitation for the convergence of the parameter estimates. Simulation case studies show the effectiveness of the proposed approach.

  • Research Article
  • Cite Count Icon 10
  • 10.1007/s00521-021-05696-2
Optimal tracking control of switched systems applied in grid-connected hybrid generation using reinforcement learning
  • Feb 4, 2021
  • Neural Computing and Applications
  • Jiayue Sun + 3 more

The paper presents a reinforcement learning approach for optimal tracking control of switched systems with application to a grid-tied hybrid generation system. To enhance interaction with the irregular environment, reference trajectory is learned via controller from states to optimal control. The main issue is to solve the optimal tracking control problem for a hybrid generation system consisting of multiple switched subsystems, and reinforcement learning can seek the globally optimal solution well without knowing accurate system dynamics. The investigated learning algorithm is used to generate an optimum map based on the learned ultimate value without knowledge of system parameters and obtains the optimal control law via deriving of algebraic Riccati equation (ARE) with unnecessary knowing of command generator dynamics. The optimal control solution can converge the online learning algorithm well based on policy iteration as verification in the simulation.

  • Research Article
  • Cite Count Icon 3
  • 10.1109/tcsi.2024.3432643
Adaptive Neural Optimal Backstepping Control of Uncertain Fractional-Order Chaotic Circuit Systems via Reinforcement Learning
  • Oct 1, 2024
  • IEEE Transactions on Circuits and Systems I: Regular Papers
  • Mei Zhong + 3 more

Optimal control has become a hot topic due to its ability to reduce control costs. However, due to the complex form of fractional-order (FO) derivatives, it is difficult to obtain the optimal control solution by solving the FO Hamilton-Jacobi-Belman equation. This article formulates an neural optimal adaptive backstepping control programme for FO chaotic circuit systems with state constraints. To avoid states exceeding constraints during optimal control, a scheme combining a transformation formula with a nonlinear state dependent function is first developed, and then the original system is transformed into an integer-order unconstrained one. To achieve optimal control, a reinforcement learning adaptive backstepping control based on the transformation scheme is introduced, where weight update laws of the reinforcement learning are constructed based on the negative gradient of a positive function rather than the square of Bellman residual, which effectively simplifies the form and design process of the update laws. According to the stability analysis, the formulated programme assures that all signals are bounded and states remain within the specified constraint space. Eventually, a simulation case is displayed to demonstrate the validity of the developed approach.

  • Supplementary Content
  • Cite Count Icon 5
  • 10.1184/r1/8397962.v1
Towards Generalization and Efficiency in Reinforcement Learning
  • Jul 2, 2019
  • Figshare
  • Wen Sun

Different from classic Supervised Learning, Reinforcement Learning (RL), is fundamentally interactive : an autonomous agent must learn howto behave in an unknown, uncertain,<br>and possibly hostile environment, by actively interacting with the environment to collect useful feedback to improve its sequential decision making ability. The RL agent will also<br>intervene in the environment: the agent makes decisions which in turn affects further evolution of the environment.<br>Because of its generality– most machine learning problems can be viewed as special cases– RL is hard. As there is no direct supervision, one central challenge in RL is how to<br>explore an unknown environment and collect useful feedback efficiently. In recent RL success stories (e.g., super-human performance on video games [Mnih et al., 2015]), we notice<br>that most of them rely on random exploration strategies, such as -greedy. Similarly, policy gradient method such as REINFORCE [Williams, 1992], perform exploration by injecting randomness into action space and hope the randomness can lead to a good sequence of actions that achieves high total reward. The theoretical RL literature has developed more sophisticated algorithms for efficient exploration (e.g., [Azar et al., 2017]), however, the sample<br>complexity of these near-optimal algorithms has to scale exponentially with respect to key parameters of underlying systems such as dimensions of state and action space. Such exponential dependence prohibits a direct application of these theoretically elegant RL algorithms to large-scale applications. In summary, without any further assumptions, RL is hard, both in practice and in theory. In this thesis, we attempt to gain purchase on the RL problem by introducing additional assumptions and sources of information.The first contribution of this thesis comes from improving RL sample complexity via imitation learning. Via leveraging expert’s demonstrations, imitation learning significantly simplifies the tasks of exploration. We consider two settings in this thesis: interactive imitation learning setting where an expert is available to query during training time, and the setting of imitation learning from observation alone, where we only have a set of demonstrations that consist of observations of the expert’s states (no expert actions are recorded). We study in both theory and in practice how one can imitate<br>experts to reduce sample complexity compared to a pure RL approach. The second contribution comes from model-free Reinforcement Learning. Specifically, we study policy<br>evaluation by building a general reduction from policy evaluation to no-regret online learning which is an active research area that has well-established theoretical foundation. Such a reduction creates a new family of algorithms for provably correct policy evaluation under<br>very weak assumptions on the generating process. We then provide a through theoretical study and empirical study of two model-free exploration strategies: exploration in action<br>space and exploration in parameter space. The third contribution of this work comes from model-based Reinforcement Learning. We provide the first exponential sample complexity separation between model-based RL and general model-free RL approaches. We then provide PAC model-based RL algorithm that can achieve sample efficiency simultaneously for many interesting MDPs such as tabular MDPs, Factored MDPs, Lipschitz continuous<br>MDPs, low rank MDPs, and Linear Quadratic Control. We also provide a more practical model-based RL framework, called Dual Policy Iteration (DPI), via integrating optimal<br>control, model learning, and imitation learning together. Furthermore, we show a general convergence analysis that extends the existing approximate policy iteration theories to DPI. DPI generalizes and provides the first theoretical foundation for recent successful practical RL algorithms such as ExIt and AlphaGo Zero [Anthony et al., 2017, Silver et al., 2017], and provides a theoretical sound and practically efficient way of unifying model-based and<br>model-free RL approaches. <br>

  • Research Article
  • Cite Count Icon 135
  • 10.1016/j.apenergy.2020.114772
Safe deep reinforcement learning-based constrained optimal control scheme for active distribution networks
  • Mar 6, 2020
  • Applied Energy
  • Peng Kou + 4 more

Safe deep reinforcement learning-based constrained optimal control scheme for active distribution networks

  • Single Report
  • 10.62311/nesx/rriv225
Optimal Control and Reinforcement Learning: Theory, Algorithms, and Robotics Applications
  • Mar 19, 2025
  • Murali Krishna Pasupuleti

Abstract: Optimal control and reinforcement learning (RL) are foundational techniques for intelligent decision-making in robotics, automation, and AI-driven control systems. This research explores the theoretical principles, computational algorithms, and real-world applications of optimal control and reinforcement learning, emphasizing their convergence for scalable and adaptive robotic automation. Key topics include dynamic programming, Hamilton-Jacobi-Bellman (HJB) equations, policy optimization, model-based RL, actor-critic methods, and deep RL architectures. The study also examines trajectory optimization, model predictive control (MPC), Lyapunov stability, and hierarchical RL for ensuring safe and robust control in complex environments. Through case studies in self-driving vehicles, autonomous drones, robotic manipulation, healthcare robotics, and multi-agent systems, this research highlights the trade-offs between model-based and model-free approaches, as well as the challenges of scalability, sample efficiency, hardware acceleration, and ethical AI deployment. The findings underscore the importance of hybrid RL-control frameworks, real-world RL training, and policy optimization techniques in advancing robotic intelligence and autonomous decision-making. Keywords: Optimal control, reinforcement learning, model-based RL, model-free RL, dynamic programming, policy optimization, Hamilton-Jacobi-Bellman equations, actor-critic methods, deep reinforcement learning, trajectory optimization, model predictive control, Lyapunov stability, hierarchical RL, multi-agent RL, robotics, self-driving cars, autonomous drones, robotic manipulation, AI-driven automation, safety in RL, hardware acceleration, sample efficiency, hybrid RL-control frameworks, scalable AI.

  • Research Article
  • Cite Count Icon 237
  • 10.1016/j.arcontrol.2012.03.004
Reinforcement learning and optimal adaptive control: An overview and implementation examples
  • Apr 1, 2012
  • Annual Reviews in Control
  • Said G Khan + 4 more

Reinforcement learning and optimal adaptive control: An overview and implementation examples

  • Research Article
  • Cite Count Icon 94
  • 10.1109/tnnls.2018.2820019
Learning-Based Predictive Control for Discrete-Time Nonlinear Systems With Stochastic Disturbances.
  • May 9, 2018
  • IEEE Transactions on Neural Networks and Learning Systems
  • Xin Xu + 3 more

In this paper, a learning-based predictive control (LPC) scheme is proposed for adaptive optimal control of discrete-time nonlinear systems under stochastic disturbances. The proposed LPC scheme is different from conventional model predictive control (MPC), which uses open-loop optimization or simplified closed-loop optimal control techniques in each horizon. In LPC, the control task in each horizon is formulated as a closed-loop nonlinear optimal control problem and a finite-horizon iterative reinforcement learning (RL) algorithm is developed to obtain the closed-loop optimal/suboptimal solutions. Therefore, in LPC, RL and adaptive dynamic programming (ADP) are used as a new class of closed-loop learning-based optimization techniques for nonlinear predictive control with stochastic disturbances. Moreover, LPC also decomposes the infinite-horizon optimal control problem in previous RL and ADP methods into a series of finite horizon problems, so that the computational costs are reduced and the learning efficiency can be improved. Convergence of the finite-horizon iterative RL algorithm in each prediction horizon and the Lyapunov stability of the closed-loop control system are proved. Moreover, by using successive policy updates between adjoint time horizons, LPC also has lower computational costs than conventional MPC which has independent optimization procedures between two different prediction horizons. Simulation results illustrate that compared with conventional nonlinear MPC as well as ADP, the proposed LPC scheme can obtain a better performance both in terms of policy optimality and computational efficiency.

  • Research Article
  • 10.24908/iqurcp18060
Reinforcement Learning for Jointly Optimal Coding and Control Policies for a Markovian System Controlled over a Communication Channel
  • Sep 9, 2024
  • Inquiry@Queen's Undergraduate Research Conference Proceedings
  • Evelyn Hubbard + 1 more

This paper develops approximation and optimality results for the optimal control of a networked system. In this system, a Markovian process is managed over a finite-rate, noiseless communication channel. Solving this problem involves determining joint optimal coding and control policies that minimize cost over time. While theoretical results on optimal coding and control structures exist, practical implementation has remained largely infeasible for non-linear systems due to the computational complexities and uncountable state spaces. This research introduces an approach that combines structural results with reinforcement learning (RL) techniques. The method uses regularity properties of the system to approximate the uncountable state space with a countable one, which allows reinforcement learning algorithms to achieve near-optimal solutions. Specifically, we establish that finite model approximations (where infinite state spaces are quantized to finite ones) and sliding finite window approximations (where a finite memory "window" of past control actions is maintained at each time step) can be employed to develop near-optimality. These approximations allow the system to be reformulated as a Markov Decision Problem (MDP) with a finite state space, making RL algorithms implementable. The convergence of the reinforcement learning algorithm to a near-optimal policy (under both the previous approximations) is supported by theoretical analysis and performance simulations. We ensure that the solutions obtained are not only computationally feasible but also nearly optimal with respect to the original problem. The applications of this work extend to many networked control systems, especially those with zero-delay coding and partially observable Markov decision processes (POMDPs). By integrating structural results with learning algorithms, this paper provides a practical framework for implementing optimal control in finite-rate environments.

  • Research Article
  • Cite Count Icon 2
  • 10.1109/tase.2025.3596555
Adaptive Backstepping Optimal Tracking Control of Interconnected Robotic Manipulator System Based on Reinforcement Learning
  • Jan 1, 2025
  • IEEE Transactions on Automation Science and Engineering
  • Hang Su + 2 more

In this paper, an optimal control method based on reinforcement learning (RL) is proposed for the tracking control problem of robotic manipulator (RM) system. In contemporary industrial manufacturing and precision technology, RM has gained significant popularity. To solve the optimal tracking control problem for RM, its dynamics are decomposed into interconnected subsystems. Combining the backstepping method and the RL method, both the virtual control signal and the actual control signal are designed as the optimal solution of each error subsystem, while ensuring that the virtual controller and the actual controller of all subsystems are optimal. For the inherent nonlinearity and difficult solution of the Hamilton–Jacobi–Bellman (HJB) equation, an actor-critic neural network (NN) is constructed to approximate the nonlinear term, and a cost function with input gain function is introduced to ensure the tracking effect. Then, based on the Lyapunov stability theory, it is proved that all the error signals of the system are semi-globally uniformly ultimately bounded (SGUUB). Finally, the effectiveness of the proposed method is validated through 2-degree-of-freedom (DOF) and 3-DOF RM simulations, with robustness verification against sudden disturbances. Note to Practitioners—The tracking control problem of RM is a very extensive and important problem in industrial applications, such as delivery, grasping, etc. The optimal control theory can solve the problems of shortest path, least fuel and shortest time. When RM perform different tasks, different cost functions can be constructed according to the control objectives. Most of the existing optimal control methods for RM systems cannot guarantee the optimality of the entire controller. Inspired by this, this paper develops a new RM optimal tracking control strategy, which can ensure that the virtual controller and the actual controller are the optimal solution, so as to ensure the overall optimality. Based on the RL method, the problem of solving the HJB equation is solved, in which the actor-critic NN weights are updated online at the same time. The actor NN is used to execute control actions, and the critic NN evaluates the control actions and feeds back to the actor NN to improve the subsequent control input.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant