Indoor Emergency Path Planning Based on the Q-Learning Optimization Algorithm

Shenghua Xu,Wenxing Jiang,Cai Chen,Yang Gu,Yu Sang,Yingyi Hu,Xiaoyan Li

doi:10.3390/ijgi11010066

Abstract

The internal structure of buildings is becoming increasingly complex. Providing a scientific and reasonable evacuation route for trapped persons in a complex indoor environment is important for reducing casualties and property losses. In emergency and disaster relief environments, indoor path planning has great uncertainty and higher safety requirements. Q-learning is a value-based reinforcement learning algorithm that can complete path planning tasks through autonomous learning without establishing mathematical models and environmental maps. Therefore, we propose an indoor emergency path planning method based on the Q-learning optimization algorithm. First, a grid environment model is established. The discount rate of the exploration factor is used to optimize the Q-learning algorithm, and the exploration factor in the ε-greedy strategy is dynamically adjusted before selecting random actions to accelerate the convergence of the Q-learning algorithm in a large-scale grid environment. An indoor emergency path planning experiment based on the Q-learning optimization algorithm was carried out using simulated data and real indoor environment data. The proposed Q-learning optimization algorithm basically converges after 500 iterative learning rounds, which is nearly 2000 rounds higher than the convergence rate of the Q-learning algorithm. The SASRA algorithm has no obvious convergence trend in 5000 iterations of learning. The results show that the proposed Q-learning optimization algorithm is superior to the SARSA algorithm and the classic Q-learning algorithm in terms of solving time and convergence speed when planning the shortest path in a grid environment. The convergence speed of the proposed Q- learning optimization algorithm is approximately five times faster than that of the classic Q- learning algorithm. The proposed Q-learning optimization algorithm in the grid environment can successfully plan the shortest path to avoid obstacle areas in a short time.

Highlights

In recent years, with the advancement of urbanization, the internal structure of urban buildings has become more complex and variable
The results show that the Q-learning optimization algorithm is better than both the SARSA algorithm and the Q-learning algorithm in terms of solving time and convergence when planning the shortest path in a grid environment
The rest of the paper is organized as below: Section 1 introduces indoor emergency path planning based on the proposed Q-learning optimization algorithm in grid environment

Summary

Introduction

With the advancement of urbanization, the internal structure of urban buildings has become more complex and variable. On the basis of Qlearning combined with ε-greedy, Li C et al [24] proposed a parameter dynamic adjustment strategy and trial-and-error action deletion mechanism, which realized the balance between adaptive adjustment and utilization in the learning process, and improved the exploration efficiency of the agent. This paper proposes a path planning algorithm based on a grid environment and optimizes the Q-learning algorithm by introducing the calculation of the exploratory factor discount rate. 2. Aimed at the problems of slow convergence speed and low accuracy of the Q-learning algorithm in a large-scale grid environment, the exploration factor in the ε-greedy strategy is dynamically adjusted, and the discount rate variable of the exploration factor is introduced. The rest of the paper is organized as below: Section 1 introduces indoor emergency path planning based on the proposed Q-learning optimization algorithm in grid environment.

Indoor Emergency Path Planning Method

Method

22.22. QQ--LLeeaarrnniinngg OOppttiimmiizzaattiioonn AAllggoorriitthhm

Path Planning Strategy

Action and state of agent

Set the reward function

Action strategy selection

Q Value Table

Dynamic Adjustment of Exploration Factors

Algorithm Flow

Algorithm Simulation Experimental Analysis

Environmental Spatial Modeling

Comparison and Analysis of Experimental Results

Simulation Scene Experiment Analysis

Experimental Data and Scene Construction

Conclusions

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: ISPRS International Journal of Geo-Information	Publication Date: Jan 14, 2022
Citations: 9	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Indoor Emergency Path Planning Based on the Q-Learning Optimization Algorithm

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: ISPRS International Journal of Geo-Information

Lead the way for us

Similar Papers

A Heuristic Indoor Path Planning Method Based on Hierarchical Indoor Modelling
Jingwen Li ... Huiqiang Wang
-
Jingwen Li, et. al.Jingwen Li ... Huiqiang Wang
01 Jan 2018
01 Jan 2018

MULTI-LEVEL INDOOR PATH PLANNING METHOD
Q Xiong ... Q Zhu
The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences | VOL. XL-4/W5
Q Xiong, et. al.Q Xiong ... Q Zhu
11 May 2015
The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences | VOL. XL-4/W5

Intelligent Vehicle Path Planning Based on Q-Learning Algorithm with Consideration of Smoothness
Wei Zhao ... Hongyan Guo
-
Wei Zhao, et. al.Wei Zhao ... Hongyan Guo
06 Nov 2020
06 Nov 2020

Research on mobile robot path planning in complex environment based on DRQN algorithm
Shuai Wang ... Yuhong Du
Physica Scripta | VOL. 99
Shuai Wang, et. al.Shuai Wang ... Yuhong Du
14 Jun 2024
Physica Scripta | VOL. 99

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Indoor Emergency Path Planning Based on the Q-Learning Optimization Algorithm

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: ISPRS International Journal of Geo-Information