DeepLoco
Learning physics-based locomotion skills is a difficult problem, leading to solutions that typically exploit prior knowledge of various forms. In this paper we aim to learn a variety of environment-aware locomotion skills with a limited amount of prior knowledge. We adopt a two-level hierarchical control framework. First, low-level controllers are learned that operate at a fine timescale and which achieve robust walking gaits that satisfy stepping-target and style objectives. Second, high-level controllers are then learned which plan at the timescale of steps by invoking desired step targets for the low-level controller. The high-level controller makes decisions directly based on high-dimensional inputs, including terrain maps or other suitable representations of the surroundings. Both levels of the control policy are trained using deep reinforcement learning. Results are demonstrated on a simulated 3D biped. Low-level controllers are learned for a variety of motion styles and demonstrate robustness with respect to force-based disturbances, terrain variations, and style interpolation. High-level controllers are demonstrated that are capable of following trails through terrains, dribbling a soccer ball towards a target location, and navigating through static or dynamic obstacles.
- Research Article
21
- 10.3390/electronics13101952
- May 16, 2024
- Electronics
The application of autonomous driving system (ADS) technology can significantly reduce potential accidents involving vulnerable road users (VRUs) due to driver error. This paper proposes a novel hierarchical deep reinforcement learning (DRL) framework for high-performance collision avoidance, which enables the automated driving agent to perform collision avoidance maneuvers while maintaining appropriate speeds and acceptable social distancing. The novelty of the DRL method proposed here is its ability to accommodate dynamic obstacle avoidance, which is necessary as pedestrians are moving dynamically in their interactions with nearby ADSs. This is an improvement over existing DRL frameworks that have only been developed and demonstrated for stationary obstacle avoidance problems. The hybrid A* path searching algorithm is first applied to calculate a pre-defined path marked by waypoints, and a low-level path-following controller is used under cases where no VRUs are detected. Upon detection of any VRUs, however, a high-level DRL collision avoidance controller is activated to prompt the vehicle to either decelerate or change its trajectory to prevent potential collisions. The CARLA simulator is used to train the proposed DRL collision avoidance controller, and virtual raw sensor data are utilized to enhance the realism of the simulations. The model-in-the-loop (MIL) methodology is utilized to assess the efficacy of the proposed DRL ADS routine. In comparison to the traditional DRL end-to-end approach, which combines high-level decision making with low-level control, the proposed hierarchical DRL agents demonstrate superior performance.
- Research Article
9
- 10.3390/app112210595
- Nov 11, 2021
- Applied Sciences
Active tracking control is essential for UAVs to perform autonomous operations in GPS-denied environments. In the active tracking task, UAVs take high-dimensional raw images as input and execute motor actions to actively follow the dynamic target. Most research focuses on three-stage methods, which entail perception first, followed by high-level decision-making based on extracted spatial information of the dynamic target, and then UAV movement control, using a low-level dynamic controller. Perception methods based on deep neural networks are powerful but require considerable effort for manual ground truth labeling. Instead, we unify the perception and decision-making stages using a high-level controller and then leverage deep reinforcement learning to learn the mapping from raw images to the high-level action commands in the V-REP-based environment, where simulation data are infinite and inexpensive. This end-to-end method also has the advantages of a small parameter size and reduced effort requirements for parameter turning in the decision-making stage. The high-level controller, which has a novel architecture, explicitly encodes the spatial and temporal features of the dynamic target. Auxiliary segmentation and motion-in-depth losses are introduced to generate denser training signals for the high-level controller’s fast and stable training. The high-level controller and a conventional low-level PID controller constitute our hierarchical active tracking control framework for the UAVs’ active tracking task. Simulation experiments show that our controller trained with several augmentation techniques sufficiently generalizes dynamic targets with random appearances and velocities, and achieves significantly better performance, compared with three-stage methods.
- Conference Article
1
- 10.1109/elecsym.2016.7861001
- Sep 1, 2016
The initial problem on building a mobile robot using embedded system is the hardware flexibility. Currently most of mobile robots use microcontroller. The weakness of building mobile robot using microcontroller is the form and the specification which is inflexible to be added or changed. For that reason this research use FPGA soft processor to design a hardware architecture in the form of system on chip (SoC) as the motor velocity control and robot heading. We divide the mobile robot control into 2 levels which are high level and low level control. High level control is including the positioning and path tracking, while low level control is including kinematic, speed control, heading lock and communication. This research focused on low level control design. The achieved result with the use of FPGA soft processor, is that the programmer can flexibly add and change the hardware architecture and faster program execution resulted from parallel process. On the motor velocity testing, the time to get back into the steady state after receiving disturbance was about 0.6 seconds. On the heading lock testing, the time of the robot to get back into the set point is around 1 second. Resources required to implement soft processor system are 3883 slices register, 1130 slices flip-flop, 3825 slices LUTs and 191 RAM Blocks.
- Research Article
30
- 10.1109/lra.2021.3092647
- Oct 1, 2021
- IEEE Robotics and Automation Letters
Modular robots have the potential for an unmatched ability to perform versatile and robust locomotion. However, designing effective and adaptive locomotion controllers for modular robots is challenging, resulting in a number of model-based methods that typically require various forms of prior knowledge. Deep reinforcement learning (DRL) provides a promising model-free approach for locomotion control by trial-and-error. However, current DRL methods often require extensive interaction data, hindering many possible applications. In this letter, a novel two-level hierarchical locomotion framework for modular quadrupedal robots is proposed. The approach combines a low-level central pattern generator (CPG)-based controller with a high-level neural network to learn a variety of locomotion tasks using DRL. The low-level CPG controller is pre-optimized to generate stable rhythmic walking gaits, while the high-level network is trained to modulate the CPG parameters for achieving task goals based on high-dimensional inputs, including the robot states and user commands. The proposed approach is employed on a simulated modular quadruped. With a limited amount of prior knowledge, the proposed method is demonstrated to be capable of learning a variety of locomotion skills such as velocity tracking, path following, and navigating to a target. Simulation results show that the proposed method can achieve higher sample efficiency than the model-free DRL method and are substantially more robust than the baseline methods to external disturbances and irregular terrain.
- Research Article
34
- 10.1016/j.ijepes.2016.10.010
- Nov 7, 2016
- International Journal of Electrical Power & Energy Systems
Simple bottom-up hierarchical control strategy for heaving wave energy converters
- Research Article
17
- 10.3390/w13060825
- Mar 17, 2021
- Water
Water loss according to water leakages in water distribution systems (WDSs) is a challenging problem worldwide. An inappropriate operation of the WDS leads to unnecessarily high pressure distribution in the WDS and thus a large amount of water leakage exists. For this reason, optimal pressure management in WDSs through regulating operations of pressure reducing valves (PRVs) is priority for water utilities. The pressure management can be accomplished in a hierarchical control scheme with high level and low level controllers. While the high level controller is responsible for calculating pressure set points for critical nodes, the task of a low level controller is to regulate the pressures at the critical nodes to the set points. The optimal pressure management in the high level controller can be casted into a nonlinear programing problem (NLP) where PRV models are crucial and determine proper operation of the WDS and quality of overall pressure control. PRV models having been used until now either describe two operating modes (active and open modes) or three operating modes (active, open and check valve modes) with parameter dependence. Such models make the formulated NLP unsuitable for the case PRVs work in check valve modes or resulted in inaccurate NLP solution with unexpected operation modes of PRVs, respectively. Therefore, this paper proposes an accurate PRV model based on complementarity constraints. The new PRV model is parameter-less dependence and is capable of describing complete operation modes of PRVs in practice. As a result, the formulated NLP is general and provides accurate NLP solution. The efficiency of our new PRV model is demonstrated on numerous case studies for optimal pressure management of WDSs.
- Research Article
65
- 10.1016/j.oceaneng.2022.110749
- Feb 9, 2022
- Ocean Engineering
COLREGs-abiding hybrid collision avoidance algorithm based on deep reinforcement learning for USVs
- Research Article
11
- 10.3390/jmse11040779
- Apr 3, 2023
- Journal of Marine Science and Engineering
The research on decision-making models of ship collision avoidance is confronted with numerous challenges. These challenges encompass inadequate consideration of complex factors, including but not limited to open water scenarios, the absence of static obstacle considerations, and insufficient attention given to avoiding collisions between manned ships and MASSs. A decision model for MASS collision avoidance is proposed to overcome these limitations by integrating the strengths of model-based and model-free methods in reinforcement learning. This model incorporates S-57 chart information, AIS data, and the Dyna framework to improve effectiveness. (1) When the MASS’s navigation task is known, a static navigation environment is built based on S-57 chart information, and the Voronoi diagram and improved A* algorithm are used to obtain the energy-saving optimal static path as the planned sea route. (2) Given the small main dimensions of an MASS, which is easily affected by wind and current factors, the motion model of an MASS is established based on the MMG model considering wind and current factors. At the same time, AIS data are used to extract the target ship (manned ship) data. (3) According to the characteristics of the actual navigation of ships at sea, the state space, action space, and reward function of the reinforcement learning algorithm are designed. The MASS collision avoidance decision model based on the Dyna-DQN model is established. Based on the DQN algorithm, the agent (MASS) and the environment interact continuously, and the actual interaction data generated are used for the iterative update of the collision avoidance strategy and the training of the environment model. Then, the environment model is used to generate a series of simulated empirical data to promote the iterative update of the strategy. Using the waters near the South China Sea as the research object for simulation verification, the navigation tasks are divided into three categories: only considering static obstacles, following the planned sea route considering static obstacles, and following the planned sea route considering both static and dynamic obstacles. The results show that through repeated simulation experiments, an MASS can complete the navigation task without colliding with static and dynamic obstacles. Therefore, the proposed method can be used in the intelligent collision avoidance module of MASSs and is an effective MASS collision avoidance method.
- Research Article
21
- 10.1038/s41598-024-81769-1
- Dec 28, 2024
- Scientific Reports
Hydrogen-based electric vehicles such as Fuel Cell Electric Vehicles (FCHEVs) play an important role in producing zero carbon emissions and in reducing the pressure from the fuel economy crisis, simultaneously. This paper aims to address the energy management design for various performance metrics, such as power tracking and system accuracy, fuel cell lifetime, battery lifetime, and reduction of transient and peak current on Polymer Electrolyte Membrane Fuel Cell (PEMFC) and Li-ion batteries. The proposed algorithm includes a combination of reinforcement learning algorithms in low-level control loops and high-level supervisory control based on fuzzy logic load sharing, which is implemented in the system under consideration. More specifically, this research paper establishes a power system model with three DC-DC converters, which includes a hierarchical energy management framework employed in a two-layer control strategy. Three loop control strategies for hybrid electric vehicles based on reinforcement learning are designed in the low-level layer control strategy. The Deep Deterministic Policy Gradient with Twin Delayed (DDPG TD3) is used with a network. Three DRL controllers are designed using the hierarchical energy optimization control architecture. The comparative results between the two strategies, Deep Reinforcement Learning and Fuzzy logic supervisory control (DRL-F) and Super-Twisting algorithm and Fuzzy logic supervisory control (STW-F) under the EUDC driving cycle indicate that the proposed model DRL-F can ensure the Root Mean Square Error (RMSE) reduction for 21.05% compared to the STW-F and the Mean Error reduction for 8.31% compared to the STW-F method. The results demonstrate a more robust, accurate and precise system alongside uncertainties and disturbances in the Energy Management System (EMS) of FCHEV based on an advanced learning method.
- Conference Article
1
- 10.1109/icm46511.2021.9385659
- Mar 7, 2021
Modular robots have the ability to perform versatile locomotion with a high diversity of morphologies. However, designing robust locomotion gaits for arbitrary robot morphologies remains exceptionally challenging. In this paper, a two-level hierarchical locomotion framework is presented for addressing modular robot locomotion tasks. The framework combines a central pattern generator controller (CPG) with a neural network trained by deep reinforcement learning. First, the low-level CPG controllers are learned by offline optimization and generate robust straight walking gaits. Second, a high-level neural network is then learned using deep reinforcement learning via trial-and-errors. The high-level learned controller can modulate the low-level CPG parameters based on online inputs including robot states and user commands. Simulation experiments are employed on a 3D modular robot. The results show that the proposed method achieves better overall performance than the baseline methods on different locomotion skills including straight walking, velocity tracking, and circular turning. Simulation results confirm the effectiveness and robustness of the proposed method.
- Conference Article
15
- 10.1115/dscc2016-9641
- Oct 12, 2016
In this paper, we present a fuel efficient control strategy for a group of connected hybrid electric vehicles (HEVs) in urban road conditions. A hierarchical control architecture is proposed where the higher level controller is located at traffic signal light while the lower level controllers are equipped on each HEV. The higher level controller utilizes Signal Phase and Timing (SPAT) information from the traffic lights to generate target velocities for every HEV, which allows a maximum number of vehicles pass the intersection at given green light window. Model Predictive Control (MPC) is used to track the target velocity and evaluate the energy efficient velocity profile for every vehicle for a given horizon. Each lower level controller then follows the velocity profile (from the higher level controller) in a fuel efficient fashion using adaptive equivalent consumption minimization strategy (A-ECMS). The lower level controller also feeds the average recuperation efficiency in a certain time window back to the higher level controller, thus affects the future velocity profile evaluation from the higher level controller, which is the major contribution of this paper. In this paper, the HEV model is developed based on Autonomie software and the simulation results show the effectiveness of our proposed approach.
- Research Article
10
- 10.3233/jifs-171186
- Apr 19, 2018
- Journal of Intelligent & Fuzzy Systems
This paper proposes an intelligent safe driving system (ISDS) for autonomous vehicles. The system utilizes a hierarchical control framework, where the high-level and low-level controllers are responsible for the decision making and motion control respectively. In the high-level controller, two finite-state machines (FSMs) are applied. One of the FSMs identifies the relative positions of the surrounding vehicles to the subject vehicle, and the other one chooses the proper driving behaviors intelligently to deal with the complex situations. In the low-level controller, the double-model-predictive-control structure is designed for the lateral motion control, and the PID feedback control with the inverse model is employed for the longitudinal motion control. The proposed control system is tested in the Simulink/CarSim simulation environment. The results show that the controlled subject vehicle is able to avoid the collision with the surrounding vehicles and acquire the desired speed autonomously. The motion stability is also guaranteed during the accelerating/decelerating and lane changing.
- Dissertation
1
- 10.32657/10356/180275
- Jan 1, 2023
Assistive wheelchairs aim to help people with mobility impairment regain mobility and provide assistance for navigation tasks. However, many prospective users, especially those with upper limb disability, find it difficult and unsafe to use a manual or powered wheelchair independently in situations where fine motor control is required, such as making sharp turns or traversing narrow doorways. The autonomous driving wheelchair fails to be their choice either because users dislike the sense of not controlling the wheelchair. Thus, the study of shared control that both helps users finish complicated navigation tasks and respects users' control authority has drawn a promising way in the field of assistive wheelchairs. Shared control approaches in assistive wheelchairs can be broadly divided into high-level shared control and low-level shared control. The high-level shared control provides different paths to the destination and predicts the user’s preferred path through the user's input. The low-level shared control executes a command for the wheelchair so that this command tries to follow the user’s instantaneous input, avoids collisions, and moves along the user's preferred path. The challenge of high-level shared control is to provide various possible paths in real-time and connect them across timesteps so that it can simultaneously predict the probabilities of all paths based on the user's control input history. Existing methods either use a pre-defined set of paths or fail to link paths across timesteps, which restricts the user's control authority. The challenge of low-level shared control is to take into account the information of various paths obtained from the high-level shared control while avoiding obstacles and following the user's instantaneous input. Most low-level shared control methods either only consider the user's instantaneous input but ignore the high-level shared control or solely assist in following the most likely path predicted by the high-level shared control. This can lead to wrong assistance when the user's input is not precise or the predicted most likely path is incorrect. In this work, a novel shared control system for point-to-point navigation of the robotic assistive wheelchair is proposed, which attempts to overcome challenges in both high-level and low-level shared control. For high-level shared control, various possible paths to the destination are computed online using the generalized Voronoi diagram and linked across time steps using their homotopy classes, and the likelihood of these paths is calculated using the principle of Bayes' rule and Maximum Entropy Inverse Optimal Control (MaxEnt IOC) based on the user input history. The low-level shared control is modeled as a Partially Observable Markov Decision Process (POMDP), a principled way of planning under uncertainty. It takes the probabilities of all paths from the high-level shared control into consideration and handles the wrong prediction through its information-gathering mechanism. User adaptability index and follow-path actions are also introduced to this POMDP model to further improve the low-level shared control. Experiments in simulations and the real wheelchair with both healthy subjects and cerebral palsy (CP) subjects prove that the assistive wheelchair with this new shared control system can significantly improve navigation outcomes for both healthy and CP subjects. The results conclude that the new shared control system reduces the time required to complete tasks and achieves better performance both subjectively and objectively.
- Research Article
62
- 10.1049/iet-its.2016.0197
- Mar 27, 2017
- IET Intelligent Transport Systems
This study presents a novel decentralised hierarchical global energy management control strategy for a group of connected four‐wheel‐drive hybrid electric vehicles (HEVs) in urban road conditions. In the higher level controller, signal phase and timing information and the optimal cruising velocity are combined to generate the target velocities for the HEVs. A model predictive control framework that focuses on the tracking of the target velocity and the associated desired control variable for every individual vehicle is proposed for the prediction of the optimal velocity that compromises fuel economy, mobility and safety. In the lower level controller, a dynamic programming problem is formulated that utilises the predicted velocity for the global energy management optimisation of every individual HEV. Simulation results validate the advantages of the proposed higher and lower level controllers.
- Conference Article
5
- 10.1109/sii.2013.6776699
- Dec 1, 2013
The time delays on the networked control system and the sensor uncertainties are the major factors affecting the performance of the autonomous mobile robot navigation in an unstructured, unknown and dynamic environment. A hierarchical, decentralized and networked control system is proposed that consists of a high-level navigation control module and a remote low-level closed-loop motion control module. The knowledge based fuzzy logic control algorithm is proposed as the high-level control. No explicit model is required. The proposed system is implemented on a two wheel differential drive mobile robot and validated through four environment scenarios as (1) slope, (2) maze, (3) dynamic obstacles and (4) garage parking. The uncertainties of the laser range finder are also evaluated. The experimental results demonstrated the proposed system can effectively and efficiently navigate the mobile robot under the effect of the time delays on the networked control system and the sensor uncertainties.