Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Towards autonomous production control: a reinforcement learning-based model for hybrid remanufacturing systems

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Growing environmental awareness and the increasing emphasis on sustainability have heightened the focus on circular economy models, particularly remanufacturing, which extends product life cycles through systematic refurbishment. Remanufacturing operations encompass three key stages: (1) remanufacturing, (2) stock management, and (3) assembly. This study optimises these interconnected processes while explicitly considering timing-related uncertainties in the remanufacturing–assembly system that impact operational performance. A remanufacturing production control model is developed to coordinate the disassembly of end-of-life products, remanufacturing of components, and hybrid assembly of final products. These finished products may integrate both remanufactured and new components, requiring effective material flow management to maintain operational efficiency and quality standards. The objective of this study is to support decision-makers in maintaining efficient production flows, meeting customer demand, and mitigating system uncertainties that may arise throughout the remanufacturing-assembly workflow. To achieve this, key performance metrics such as throughput, machine-related disruptions, and material release are continuously monitored to guide system optimisation. A reinforcement learning (RL)-based approach is proposed to determine the optimal material flow, service level management, and on-time delivery of remanufactured products within a hybrid assembly environment. The assembly process operates under three distinct operational modes, allowing the RL agent to dynamically adapt its decision-making strategy based on real-time system conditions. Additionally, the agent explicitly accounts for machine failures in its decision framework to proactively manage disruptions and maintain a stable material flow. Simulation results demonstrate that the decentralised RL-based production control system outperforms conventional heuristics, achieving an average cumulative reward improvement of approximately 4% compared to the best-performing combination of Constant Work-In-Progress (CONWIP) and the One-Disassembly method, and 33% compared to the combination of CONWIP and the Simple Disassembly method. The findings suggest that RL is suitable for addressing typical timing variability and disruption effects in remanufacturing systems, including stochastic return timing, condition-dependent processing-time variability, and machine failures.

Similar Papers
  • Research Article
  • 10.17309/tmfv.2025.3.04
Defining the Influence of Age and Gender on Key Performance Metrics in Badminton
  • May 30, 2025
  • Physical Education Theory and Methodology
  • Titis Pambudi + 5 more

Background. Badminton, a racquet sport that has gained global popularity, demands technical precision, tactical awareness, and exceptional physical fitness. Skills such as smashing, footwork, and minimizing errors are critical to success. However, the specific influence of age and gender on these metrics, especially among younger players, remains underexplored. Objectives. This study aimed to examine the effect of age and gender on key badminton performance metrics, including smash ability, footwork, and unforced errors, in order to identify developmental and demographic factors influencing skill acquisition and execution. Materials and methods. A quantitative descriptive study involved 24 athletes (aged 9–14) from the Wincorp badminton organization in Surakarta, Indonesia. Participants were grouped by age (9–10, 11–12, 13–14 years) and gender, ensuring equal representation. Over two months, data on smashing, lobbing, driving, footwork, and error rates were collected. Descriptive statistics and MANOVA analyzed differences, with a significance level set at p < 0.05. Results. MANOVA revealed significant age-related effects on smashing (p = 0.000), footwork (p = 0.000), and error points (p = 0.000), with beginners (13–14 years) excelling in most metrics. Gender differences were also found to be substantial for smashing (p = 0.000), footwork (p = 0.000), and error points (p = 0.003), with males outperforming females in most categories. Interaction effects between age and gender were significant for smashing and footwork (p < 0.05). However, no considerable differences were observed for netting and serving strokes across age or gender. Conclusions. The study indicates that age and gender significantly influence badminton performance metrics. Beginner athletes (13–14 years) demonstrated superior skills compared to younger groups, while males generally outperformed females. These findings highlight the importance of tailoring training programs by age and gender to optimize skill development and reduce performance gaps. Further studies should be performed to investigate biomechanical and psychological factors to refine coaching strategies.

  • Research Article
  • 10.1080/21681724.2025.2602000
Enhanced Ion/Ioff ratio and RF performance for gate all around (GAA) AlGaN/GaN with GaN cap layer fin-HEMT for higher gate capacitance for improved RF communication
  • Dec 17, 2025
  • International Journal of Electronics Letters
  • Lakhmikanta Mishra + 2 more

In this study, we conducted a comprehensive DC and AC analysis of Gate-All-Around (GAA) devices, focusing on the implications of increased gate capacitance both for wider and narrower device on key performance metrics. Our findings reveal that the elevated gate capacitance due to thin GaN cap layer significantly affects the transconductance (gm), Drain-Induced Barrier Lowering (DIBL), subthreshold swing (SS) and the on/off current ratio (Ion/Ioff). A new analytical linkage is formulated connecting gate capacitance ( C g ) with key DC performance metrics. The results indicated a significant enhancement in both ft and fmax, attributed to the high Ion/Ioff ratio of 1.72 × 10 13 , confirming the superior high-frequency performance of GAA structures. Our analysis underscores the critical role of gate capacitance in optimising the electrical characteristics of GAA devices, paving the way for their application in 5G wireless communication.

  • Research Article
  • 10.3389/fmed.2025.1623776
Decentralized community-integrated research sites drive higher randomization rates: insights from a large-scale neurodegenerative disease trial
  • Aug 19, 2025
  • Frontiers in Medicine
  • Seyed-Mahdi Khaligh-Razavi + 3 more

IntroductionRecruitment and retention remain critical challenges in clinical trials, particularly in neurodegenerative diseases, which require large participant populations, rigorous screening, and prolonged follow-up periods. Care Access is a global research site management organization that operates clinical trial sites employing various operational models. This study evaluates the operational performance of Care Access site models—including traditional sites, hub-and-spoke, and decentralized community-integrated research (DCIR) sites—within a Phase 3 neurodegenerative disease trial, focusing on their relative efficiency in recruitment, randomization, and retention. The inclusion of multiple site models within the same trial presents a rare opportunity for direct comparison under uniform study conditions, providing unique insights into their respective advantages and challenges. By analyzing key site performance metrics and the role of innovative operational strategies, this study aims to identify effective approaches to enhancing trial efficiency and overcoming recruitment challenges to inform the design and conduct of future trials.MethodsThe trial involved 32 Care Access sites each employing one of these distinct operational models. Key performance metrics, such as participant screening rates, randomization rates, screen failure rates, and post-randomization discontinuation rates, were analyzed across (a) traditional, (b) hub-and-spoke, and (c) DCIR site models. We also compared the enrollment performance of Care Access to that of 196 non-Care Access sites using publicly available data.ResultsDCIR Sites demonstrated the highest recruitment efficiency, screening 20.61 participants per site per month and randomizing 0.79 participants per site per month, compared to 11.78 and 0.50 for traditional sites, and 12.20 and 0.45 for hub-and-spoke sites, respectively. Despite being newly established, and operating in a decentralized model, DCIR sites achieved post-randomization discontinuation rates (28.17%) comparable to those of traditional site models (26.28%), highlighting their effectiveness in maintaining participant engagement. All site models encountered high screen failure rates (~95%), consistent with Phase 3 trials for neurodegenerative diseases. Notably, a community-engaged, research-only facility achieved the lowest discontinuation rate (17.65%) among all sites, highlighting the potential of strong local engagement to significantly enhance retention and participation. Furthermore, when comparing Care Access sites with non-Care Access sites in this trial, Care Access sites achieved an average randomization rate of 15.6 participants per site, outperforming the 8.7 participants per site recorded by non-Care Access sites. Data quality, monitoring practices, and overall data integrity were consistent across all site models, supporting the reliability of findings across both decentralized and traditional approaches. This comparison highlights the effectiveness of the innovative operational framework and decentralized community engagement approach in overcoming traditional recruitment challenges and enhancing trial outcomes.DiscussionDCIR sites exhibited superior participant screening and randomization efficiency while maintaining discontinuation rates comparable to traditional site models. This success was driven by a combination of innovative operational strategies, including decentralized community-based outreach mechanisms that expanded population access to research by bringing trials directly to populations that previously lacked access to clinical research. At the same time, this approach helped reach underrepresented groups, thereby improving both geographic coverage and trial generalizability while enhancing overall trial performance. Additionally, other innovations like the deployment of centralized remote research coordinators also played a role by streamlining remotely-conducted tasks, allowing site staff, in all site models, to focus on participant care and engagement. These findings highlight the effectiveness of a flexible, multi-model site strategy in addressing recruitment and retention challenges in large-scale Phase 3 neurodegenerative disease trials and suggest that this approach may extend to other therapeutic areas facing similar challenges.

  • Research Article
  • 10.5465/ambpp.2019.12973abstract
I Got 1099 Problems but Finding a Ride Ain’t One: Conflict Resolution in the Ridehail Industry
  • Aug 1, 2019
  • Academy of Management Proceedings
  • Michael Maffie

In 2017, Uber announced it had a “broken relationship” with its drivers. Yet what broke the relationship and what effect might it have on Uber’s organizational performance? This paper is a mixed-methods study of how conflict impacts the relationship between platforms and drivers. First, by drawing on 55 original interviews with rideshare drivers, this paper identifies conflict triggers that damage the relationship between platforms and drivers. Second, this paper uses new time diary data from 490 Uber, Lyft, Juno, and other Transportation Network Company (TNC) drivers from across the United States to empirically test if these conflict triggers are associated with key TNC performance metrics. Specifically, this paper finds that a “broken relationship” is associated with drivers withholding their working time from Uber and allocating it toward their competitor, Lyft. Additionally, this paper finds that drivers who have a broken relationship with Uber are more likely to recruit (‘steer’) passengers toward Lyft. These behaviors provide a link between organizational conflict and key performance metrics in the ‘gig economy.’

  • Conference Article
  • Cite Count Icon 6
  • 10.1117/12.672448
Optical verification of the James Webb Space Telescope
  • Jun 14, 2006
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Brian Mccomas + 6 more

The optical system of the James Webb Space Telescope (JWS T) is split between two of th e Observatory’s element, the Optical Telescope Element (OTE) and the Integrated Science Instrument Module (ISIM). The OTE optical design consists of an 18-hexagonal segmented primary mirror (25m 2 clear aperture), a secondary mirror, a tertiary mirror, and a flat fine steering mirror used for fine guidance control. All optical components are made of beryllium. The primary and secondary mirror elements have hexapod actuation that provides six degrees of freedom rigid body adjustment. The optical components are mounted to a very stable truss struct ure made of composite materials. The OTE structure also supports the ISIM. The ISIM contains the Science Instruments (SIs) and Fine Guidance Sensor (FGS) needed for acquiring mission science data and for Observatory pointing and control and provides mechanical support for the SIs and FGS. The optical perform ance of the telescope is a key performance me tric for the success of JWST. To ensure proper performance, the JWST optical verification program is a comprehensive, incremental, end-to-end verification program which includes multiple, independent, cross checks of key optical performance metrics to reduce risk of an on-orbit telescope performance issues. This paper discusses the verification testing and analysis necessary to verify the Observatory’s image quality and sensitivity requirements. This verification starts with component level verification and ends with the Observatory level verification at Johnson Space Flight Center. The optical verification of JWST is a comprehensive, incremental, end-to-end optical veri fication program which includes both test and analysis. Keywords: JWST, Optical Verification, Cryogenic Te sting, Optical Analysis, Optical Testing

  • Research Article
  • 10.3390/electronics14152982
Impact of Charge Carrier Trapping at the Ge/Si Interface on Charge Transport in Ge-on-Si Photodetectors
  • Jul 26, 2025
  • Electronics
  • Dongyan Zhao + 8 more

The performance of optoelectronic devices is affected by various noise sources. A notable factor is the 4.2% lattice mismatch at the Ge/Si interface, which significantly influences the efficiency of Ge-on-Si photodetectors. These noise sources can be analyzed by examining the impact of the Ge/Si interface and deep traps on dark and photocurrents. This study evaluates the impact of these charge traps on key photodetector performance metrics, including responsivity, photo-to-dark current ratio, noise equivalent power (NEP), and specific detectivity (D*). The trapping effects on charge transport under both forward and reverse bias conditions are monitored through hysteresis analysis. When illuminated with an unmodulated 1550 nm laser, all the key performance metrics exhibit maximum variations at a specific reverse bias. This critical bias marks the transition from saturated to exponential charge transport regimes, where intensified electric fields enhance trap-assisted recombination and thus maximize metric fluctuations.

  • Conference Article
  • Cite Count Icon 10
  • 10.1109/icastech.2014.7068061
Performance Analysis of R-DCN Architecture for Next Generation Web Application Integration
  • Oct 1, 2014
  • C C Udeze + 4 more

With the astronomical growth in online presence vis-a-vis the IT industry across the globe, there is an urgent need to evolve cloud based DataCenter architectures that can rapidly accommodate web application integrations which can serve the energy industries, educational sector, finance sector, manufacturing sector, etc. Previous works in DataCenter domain have not investigated on cloud based DataCenter Servercentric characteristics for evaluation studies. This paper then proposed a Reengineered DataCenter (R-DCN) architecture for efficient web application integration. In this regard, attempt was made to address network scalability, efficient and distributed routing, packets traffic control, and network security using analytical formulations and behavioural algorithms having considered some selected architectures. In this work a simulation experiment was carried out to study six key performance metrics for all the architectures. It was observed that the network throughput, fault-tolerance/network capacity, utilization, latency, service availability, scalability and clustering effects of R-DCN responses with respect to above Key Performance (KPIs) metrics were satisfactory when compared with BCube and DCell architectures. Future work will show a detailed validation using a cloud testbed and CloudSim Simulator.

  • Research Article
  • 10.19101/ijatee.2024.111100094
Quantum prioritized experience replay with MaDi-based priority and quantum circuit mechanisms for optimizing reinforcement learning
  • Dec 31, 2024
  • International Journal of Advanced Technology and Engineering Exploration
  • R Palanivel + 1 more

Reinforcement learning (RL) encounters significant challenges related to scalability and computational efficiency, particularly in complex decision-making environments. Traditional RL algorithms often struggle with large-scale tasks, necessitating innovative approaches to enhance adaptability and performance. The quantum circuit-based priority replay (QCPR) algorithm was introduced in this study, leveraging quantum-inspired techniques to enhance learning efficiency and decision-making quality through quantum computing (QC). The QCPR algorithm implements two key mechanisms for prioritizing experiences: the magnitude and direction (MaDi) based priority, which assesses the significance and directionality of experiences, and QCPR-specific methods that ensure efficient experience replay. These mechanisms are seamlessly integrated into the Q-learning (QL) framework to optimize the RL process for superior performance. QCPR was evaluated against four well-established RL algorithms-QL, deep Q-networks (DQN), prioritized experience replay (PER), and dueling DQN (DDQN)-using standard decision-making benchmarks. The algorithm was further tested in various simulation and real-world environments, including Atari games, Qiskit integration, and Raspberry Pi hardware and software setups. These evaluations demonstrated QCPR’s adaptability and robustness, showcasing its capability for dynamic, large-scale applications. The results revealed that QCPR significantly outperforms its counterparts across key performance metrics, including accuracy, precision, recall, F1-score, and cumulative rewards. Specifically, QCPR achieved a 15.29% improvement in accuracy over QL, a 9.98% increase over DQN, a 6.33% improvement over PER, and an 8.23% enhancement over DDQN. This study highlights the potential of quantum-inspired approaches to advance RL, offering scalable and efficient solutions for complex decision-making tasks.

  • Research Article
  • Cite Count Icon 1
  • 10.1080/00295450.2025.2515654
Reinforcement Learning for Anomaly Detection in Nuclear Power Plant Operation and Maintenance
  • Aug 1, 2025
  • Nuclear Technology
  • Ezgi Gursel + 6 more

In nuclear power plants (NPPs), timely identification of sensor and human errors is critical to ensure safe and efficient plant operations. Anomaly detection models can be employed for this task. However, traditional anomaly detection approaches may have high dependency on labeled data sets and struggle with adaptability in complex, dynamic environments. Reinforcement learning (RL) has demonstrated significant potential in fault diagnosis and anomaly detection; however, its application to anomaly detection in NPPs remains a relatively underexplored research direction. Hence, to address this gap, in this study, we present a novel physics-informed RL model, PIRL-AD (physics-informed reinforcement learning for anomaly detection), that integrates domain knowledge from calorimetric equations into the RL framework for enhanced sensor and human error anomaly detection. We evaluate the performance of PIRL-AD against a nonphysics-informed RL benchmark and a support vector machine (SVM) on data collected from a forced-flow loop testbed. Experimental results suggest that PIRL-AD outperforms other baselines on a range of anomalous data sets that include both sensor and human-induced anomalies across key performance metrics, statistically outperforming the RL and SVM benchmarks with respect to geometric mean (respectively, 92.96% versus 91.06% versus 83.01%) and F1 score (respectively, 89.23% versus 86.98% versus 77.01%). The findings suggest the potential of physics-integrated RL models for enhanced anomaly detection performance in NPPs.

  • Research Article
  • Cite Count Icon 2
  • 10.7282/t3xg9rmw
Performance analysis and design of batch ordering policies in supply chains
  • Jan 1, 2007
  • Rutgers University Community Repository (Rutgers University)
  • Abdullah S Karaman

Devising manufacturing/distribution strategies for supply chains and determining their parameter values have been challenging problems. Linking production management to stock keeping processes improves the planning of the supply chain activities, including material management, culminating in improved customer service levels. In this thesis, we investigate a multi-echelon supply chain consisting of a supplier, a plant, a distribution center and a retailer. Material flow between stages is driven by reorder point/order quantity inventory control policies. We develop a model to analyze supply chain behavior using some key performance metrics such as the time averages of inventory and backorder levels, as well as customer service levels at each echelon. The model is validated via simulation, yielding good agreement of robust performance metrics. The metrics are then used within an optimization framework to help design the supply chain by calculating optimal parameter values minimizing the expected total cost. Optimal design of the material flow system is part of the overall planning and operation of a supply chain. The outcome of the optimization framework specifies not only how much and where to hold inventory but also how to move inventory across the supply chain. The developed model requires limited computational requirements, which in turn helps frequently update the performance measures and optimal system parameters so as to be more responsive to short-term changes in demand or supply. In addition, it can be used as a decision support system for effective decision making as opposed to using simplistic inventory models, which results in significantly higher operating costs. In a similar vein, we consider a distribution inventory system with one warehouse and several retailers. The challenge in this system is to describe the demand arrival process at the warehouse. We propose a procedure to characterize the demand arrival process at the warehouse as a superposition of several independent Erlang processes. An important characteristic of the superposed process is that although the individual processes are independent from each other, the superposed process may be no longer independent. We present a methodology to characterize such arrival streams as Markovian processes. We, then, extend the methodology to phase-type arrival streams as well.

  • Research Article
  • Cite Count Icon 1
  • 10.3390/electronics14081494
Location-Based Handover with Particle Filter and Reinforcement Learning (LBH-PRL) for Mobility and Service Continuity in Non-Terrestrial Networks (NTN)
  • Apr 8, 2025
  • Electronics
  • Li-Sheng Chen + 2 more

In high-mobility non-terrestrial networks (NTN), the reference signal received power (RSRP)-based handover (RBH) mechanism is often unsuitable due to its limitations in handling dynamic satellite movements. RSRP, a key metric in cellular networks, measures the received power of reference signals from a base station or satellite and is widely used for handover decision-making. However, in NTN environments, the high mobility of satellites causes frequent RSRP fluctuations, making RBH ineffective in managing handovers, often leading to excessive ping-pong handovers and a high handover failure rate. To address this challenge, we propose an innovative approach called location-based handover with particle filter and reinforcement learning (LBH-PRL). This approach integrates a particle filter to estimate the distance between user equipment (UE) and NTN satellites, combined with reinforcement learning (RL), to dynamically adjust hysteresis, time-to-trigger (TTT), and handover decisions to better adapt to the mobility characteristics of NTN. Unlike the location-based handover (LBH) approach, LBH-PRL introduces adaptive parameter tuning based on environmental dynamics, significantly improving handover decision-making robustness and adaptability, thereby reducing unnecessary handovers. Simulation results demonstrate that the proposed LBH-PRL approach significantly outperforms conventional LBH and RBH mechanisms in key performance metrics, including reducing the average number of handovers, lowering the ping-pong rate, and minimizing the handover failure rate. These improvements highlight the effectiveness of LBH-PRL in enhancing handover efficiency and service continuity in NTN environments, providing a robust solution for intelligent mobility management in high-mobility NTN scenarios.

  • Research Article
  • Cite Count Icon 2
  • 10.48175/ijarsct-18785
Reinforcement Learning for Adaptive Cognitive Sensor Networks
  • Jun 8, 2024
  • International Journal of Advanced Research in Science, Communication and Technology
  • Nazeer Shaik + 1 more

In this paper, we propose an adaptive cognitive sensor network (CSN) system utilizing reinforcement learning (RL) to optimize network performance dynamically. The RL-based system adjusts key parameters such as transmission power, channel selection, and data scheduling based on real-time environmental feedback, thereby enhancing energy efficiency, spectrum utilization, and data accuracy. A Q-learning algorithm is employed to train the RL agent, which operates under an ϵ-greedy policy to balance exploration and exploitation. Comparative analysis with traditional static and rule-based systems demonstrates significant improvements across all key performance metrics. Future enhancements are suggested, including advanced RL techniques, transfer learning, and real-world deployments, highlighting the potential of RL in transforming CSNs into more intelligent, efficient, and resilient networks

  • Conference Article
  • Cite Count Icon 1
  • 10.1115/omae2022-78184
Analysis of Life Extension Performance Metrics for Offshore Wind Assets
  • Jun 5, 2022
  • Baran Yeter + 2 more

The objective of the present study is to investigate systematically the key metrics to evaluate the life extension performance of offshore wind farm operations. Finding the appropriate performance metric for an operation is essential for a durable, reliable, and profitable offshore wind farm operation. The analyzed key performance metrics are gross profit margin, return on asset, compounded annual rate of return of initial investment and levelized cost of energy. The mean value and standard deviation of each performance metric are calculated within a probabilistic techno-economic assessment framework for a single offshore wind asset, which is later extended to evaluate the whole offshore wind farm by a multi-asset portfolio optimization. The Markowitz modern portfolio theory is applied to estimate the maximum risk-adjusted ratio and Sharpe ratio, for the key performance metrics. Subsequently, the key performance metrics are compared to identify the most suitable metrics at different stages of the life extension. Moreover, the present study investigates the effect of different uncertainty levels associated with the stochastic variables in the techno-economic assessment. Finally, the suitability of performance metrics is analyzed and discussed for different offshore wind farm sizes and related recommendations are given.

  • Research Article
  • Cite Count Icon 37
  • 10.1108/mabr-03-2020-0018
Revising the warehouse productivity measurement indicators: ratio-based benchmark
  • Aug 11, 2020
  • Maritime Business Review
  • Nur Hazwani Karim + 6 more

Purpose The literature on warehouse performance assessments is mainly focussed on the efficiency and effectiveness of an action or activity due to customer demand and tailored fulfilment, with less attention being given to the performance measurement of each function of the warehouse and its overall productivity. Therefore, this study was aimed at revising the key warehouse performance metrics to a set of productivity measurement indicators that can be adopted internationally for benchmarking productivity performance. Design/methodology/approach A literature review and semi-structured survey questionnaire were used for this study. The importance of warehouse productivity performance was reviewed to revamp the measurement indicators. Through the use of a directed content analysis and descriptive analysis, an extensive study was carried out to analyze existing warehouse productivity indicators. Findings The findings of this study provide comprehensive references for practitioners and academicians for improving the classification of productivity measurements from existing key performance metrics for warehousing. Also, this paper highlights the warehouse resources related to the respective warehouse operation activities. Research limitations/implications The study was limited to productivity performance indicators adapted from Staudt et al. (2015). Furthermore, the samples for this study comprised Malaysian academicians and practitioners in the related field. The findings can be adapted on a global scale as this study implemented general warehouse operation processes. Originality/value Consequently, the contributions of this study are that it provides relevant benchmarks for key productivity performance indicators in the warehousing sector that has worldwide applicability and the developed model provides a conceptual platform from which further theoretical and empirical developments can be carried out.

  • Research Article
  • 10.1038/s41598-026-51502-1
Goal-guided greedy experience replay-enhanced reinforcement learning for efficient autonomous navigation.
  • May 19, 2026
  • Scientific reports
  • Yichun Zeng + 1 more

Despite some success in mapless goal-driven navigation using deep reinforcement learning, there is an issue of insufficient experience utilization in deep reinforcement learning-based mapless goal-driven navigation. The reason is that during experience sampling, the differences between experiences are not fully considered. Uniform sampling leads to the underutilization of experiences that are more beneficial for agent learning. Current deep reinforcement learning-based mapless goal-driven navigation approaches fail to adequately account for this, resulting in low experience utilization efficiency during the agent's learning process. To address this issue, we propose a Goal-guided Greedy Experience Replay Enhanced Reinforcement Learning (GER-RL) method for efficient autonomous navigation. More specifically, we prioritize experiences by their importance, employing non-uniform sampling, and incorporate this experience sampling approach into the reinforcement learning-based navigation model to improve data utilization efficiency. Experiments conducted in a simulation environment show that our method prioritizes experiences that are more useful to the agent, improving the data efficiency of the DRL learning process and significantly enhancing navigation performance. Compared to existing deep reinforcement learning-based methods for mapless goal-driven navigation, our approach demonstrates significant improvements across key performance metrics, including average reward, average length, average time, success rate and collision rate.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant