Articles published on Reinforcement learning
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
58869 Search results
Sort by Recency
- New
- Research Article
- 10.1016/j.ijmedinf.2026.106426
- Jul 1, 2026
- International journal of medical informatics
- Yahya Kemal Çalışkan + 3 more
Beyond block time: a head-to-head comparison of reinforcement learning, genetic algorithms, and predict-then-optimize scheduling for operating room workflow using discrete-event simulation.
- New
- Research Article
- 10.1016/j.watres.2026.125855
- Jul 1, 2026
- Water research
- Jeongwoo Moon + 6 more
Adaptive reinforcement learning for energy-efficient high-recovery closed-circuit reverse osmosis.
- New
- Research Article
- 10.1016/j.jbi.2026.105031
- Jul 1, 2026
- Journal of biomedical informatics
- Pengze Li + 12 more
Exploring the role of reinforcement learning in vision-language models for cardiovascular disease decision support.
- New
- Research Article
- 10.1016/j.neunet.2026.108667
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Gong Gao + 3 more
Existing value-based online reinforcement learning (RL) algorithms suffer from slow policy exploitation due to ineffective exploration and delayed policy updates. To address these challenges, we propose an algorithm called Instant Retrospect Action (IRA). Specifically, we propose Q-Representation Discrepancy Evolution (RDE) to facilitate Q-network representation learning, enabling discriminative representations for neighboring state-action pairs. In addition, we adopt an explicit method to policy constraints by enabling Greedy Action Guidance (GAG). This is achieved through backtracking historical actions, which effectively enhances the policy update process. Our proposed method relies on providing the learning algorithm with accurate k-nearest-neighbor action value estimates and learning to design a fast-adaptable policy through policy constraints. We further propose the Instant Policy Update (IPU) mechanism, which enhances policy exploitation by systematically increasing the frequency of policy updates. We further discover that the early-stage training conservatism of the IRA method can alleviate the overestimation bias problem in value-based RL. Experimental results show that IRA can significantly improve the learning efficiency and final performance of online RL algorithms on eight MuJoCo continuous control tasks.
- New
- Research Article
- 10.1016/j.artmed.2026.103413
- Jul 1, 2026
- Artificial intelligence in medicine
- Kenneth Lau + 7 more
Reinforcement learning for real-time adaptive radiotherapy.
- New
- Research Article
- 10.1016/j.engappai.2026.114598
- Jul 1, 2026
- Engineering Applications of Artificial Intelligence
- Xiaokang Ma + 4 more
Real-time adaptive energy management strategy based on multi-agent collaborative decision-time planning for off-road hybrid electric vehicles
- New
- Research Article
- 10.1016/j.conengprac.2026.106890
- Jul 1, 2026
- Control Engineering Practice
- V Burgaud + 3 more
Reinforcement learning vs. model-based control in electric vehicle charging microgrids
- New
- Research Article
- 10.1016/j.asoc.2026.115168
- Jul 1, 2026
- Applied Soft Computing
- Li Long + 7 more
Algorithmic trading by reinforcement learning in a collaborative manner
- New
- Research Article
- 10.1016/j.neunet.2026.108735
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Tenglong Yang + 4 more
Adaptive exploration strategy in reinforcement learning based on Q-values and environmental cognition.
- New
- Research Article
- 10.1016/j.engappai.2026.114694
- Jul 1, 2026
- Engineering Applications of Artificial Intelligence
- Xiaoning Shen + 3 more
Agile software project scheduling using reinforcement learning with genetic programming
- New
- Research Article
- 10.1002/mp.70546
- Jul 1, 2026
- Medical physics
- Christopher Huynh + 5 more
Inverse planning is often used for Gamma Knife radiosurgery, allowing clinicians to mathematically specify desired clinical objectives and dose limits. The objectives are controlled by weights that are manually tuned to find the desired trade-off, which varies from case to case. Automation of this process can reduce clinical workload and improve consistency in plan quality. To train a deep reinforcement learning agent using a reward function that incorporates the clinical metrics from past plans into its scoring criteria. The metric trade-off from the clinical plan is scored higher than all others, guiding the agent to produce plans with similar trade-offs. An agent was trained to adjust the two priority weights (i.e., digital slider bars) in the clinical inverse planner. The agent consists of a neural network that receives the metrics and dose distribution of the current plan and the target and organ-at-risk masks as inputs. These methods were demonstrated on a dataset of 204 single-target metastases and a dataset of 71 acoustic neuroma cases. The cases were split into training, validation, and testing sets of size 123/41/40 and 42/14/15 for the metastases and acoustic neuromas, respectively. On the metastases test dataset, the agent achieved a significantly higher (p=0.0136) average plan score (3.925±0.130) compared to the default slider plans (3.874±0.147). On the acoustic neuromas test dataset, the agent achieved a higher (p=0.4493) average plan score (4.035±0.177) compared to the default slider plans (3.995±0.365). The higher plan scores are reflected in the four plan quality metrics: the agent's plans, on average, had metrics more similar to the clinical plans, compared to the default slider plans, for both test datasets. The proposed reward function enabled the agent to learn to find plans that aligned with historical planning decisions. Future work will investigate providing the agent with additional inputs that can explain the variability in planning decisions, which would further improve its performance.
- New
- Research Article
- 10.1016/j.asoc.2026.115192
- Jul 1, 2026
- Applied Soft Computing
- Zhensong Chen + 4 more
Adversarial fraud sample generation with reinforcement learning: A joint optimization framework for financial statement fraud detection
- New
- Research Article
- 10.1016/j.chaos.2026.118248
- Jul 1, 2026
- Chaos, Solitons & Fractals
- Yanfen Song + 3 more
Reinforcement learning and game-based optimal output consensus control for higher-order multi-agent systems with unknown dead-zone inputs
- New
- Research Article
- 10.1016/j.neunet.2026.108644
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Yue Zhou + 3 more
Observer-based prescribed-time optimal neural consensus control for six-rotor UAVs: A novel actor-critic reinforcement learning strategy.
- New
- Research Article
- 10.1097/yco.0000000000001091
- Jul 1, 2026
- Current opinion in psychiatry
- Germano Vera Cruz + 2 more
Addictive behaviors, including both substance use disorders and behavioral addictions, arise from complex interactions among biological, psychological, social, and environmental factors including digital ones. This review focuses on the assessment of social and psychological risk and protective factors, highlighting how artificial intelligence and machine learning approaches complement conventional qualitative and quantitative methodologies. The aim is to clarify how these tools can enhance understanding, prediction, and prevention of addictive behaviors. Recent research identifies impulsivity, emotion dysregulation, peer norms, and family functioning as central psychosocial risk factors for addictive behaviors. Protective factors - such as self-efficacy, social support, and family cohesion - moderate these risks. Conventional analyses provide foundational evidence, while ML methods (predictive machine learning, explainable artificial intelligence, reinforcement learning) now enable integration of multimodal data, detection of nonlinear patterns, and identification of latent psychosocial profiles. Emerging studies demonstrate potential for early-warning prediction and personalized intervention design. AI/ML offers unprecedented opportunities to advance addiction science by handling high-dimensional psychosocial and behavioral data. Yet, ethical, interpretative, and causal challenges persist. The most promising path forward lies in synergizing theory-driven analytics with data-driven AI approaches to achieve more precise and contextually grounded prevention and intervention strategies for addictive behaviors.
- New
- Research Article
- 10.1016/j.eswa.2026.131850
- Jul 1, 2026
- Expert Systems with Applications
- Jiahao Pan + 2 more
Efficient and safe decision-making in reinforcement learning: One-step anticipatory policy selector with adaptive safety thresholds
- New
- Research Article
- 10.1016/j.cct.2026.108354
- Jul 1, 2026
- Contemporary clinical trials
- Rachel T Gonzalez + 5 more
Practical considerations when designing an online learning algorithm for an app-based mHealth intervention.
- New
- Research Article
- 10.1016/j.neunet.2026.108693
- Jul 1, 2026
- Neural networks : the official journal of the International Neural Network Society
- Fandi Gou + 2 more
A graph-based safe reinforcement learning method for multi-agent cooperation.
- New
- Research Article
- 10.1016/j.cor.2026.107444
- Jul 1, 2026
- Computers & Operations Research
- Mariusz Kaleta + 1 more
A neural-driven constructive heuristic for the flexible job shop scheduling problem: An efficient alternative to complex deep learning methods
- New
- Research Article
- 10.1109/tcyb.2026.3660478
- Jul 1, 2026
- IEEE transactions on cybernetics
- Qi Duan + 3 more
This study develops a reinforcement learning (RL)-based control framework with guaranteed predefined performance for nonlinear switched interconnected systems. This approach effectively addresses challenges arising from unmeasurable states and group average dwell time switching mechanisms, allowing both convergence time and accuracy to be preset via parameter configuration. First, the system equations are reconstructed to target nonlinear and interconnected terms, which are then approximated using neural networks (NNs). Additionally, an NNs-based switching state observer is designed to estimate the unmeasurable states. Second, within the backstepping synthesis framework, a distributed optimal controller is designed by integrating a performance transformation function into the cost function, with the resulting control law approximated via an identifier-actor-critic architecture. Furthermore, the group average dwell time-based stability analysis is generalized to address the optimal control challenges inherent in nonlinear switched interconnected systems. Compared with existing studies, this approach demonstrates enhanced extensibility and practicality for real-world applications. Finally, two simulation examples verify the effectiveness and superiority of the proposed method over state-of-the-art alternatives.