Adam: A Method for Stochastic Optimization
Adam is an efficient, easy-to-implement stochastic optimization algorithm utilizing adaptive estimates of lower-order moments, suitable for large-scale, non-stationary, and noisy problems. It demonstrates favorable empirical performance and theoretical convergence guarantees, with a variant called AdaMax based on the infinity norm.
We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments. The method is straightforward to implement, is computationally efficient, has little memory requirements, is invariant to diagonal rescaling of the gradients, and is well suited for problems that are large in terms of data and/or parameters. The method is also appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients. The hyper-parameters have intuitive interpretations and typically require little tuning. Some connections to related algorithms, on which Adam was inspired, are discussed. We also analyze the theoretical convergence properties of the algorithm and provide a regret bound on the convergence rate that is comparable to the best known results under the online convex optimization framework. Empirical results demonstrate that Adam works well in practice and compares favorably to other stochastic optimization methods. Finally, we discuss AdaMax, a variant of Adam based on the infinity norm.
- Book Chapter
2
- 10.5772/14951
- Feb 28, 2011
Global and dynamic optimization of engineering problems usually involves complex physico-chemical models as constraints. These models are in general highly non-linear, resulting in multimodal optimization problems. The model may have discontinuous behavior and/or include a very large set of variables. As the complexity of the systems increases, equation-free modeling is becoming more common (Kevrekidis, Gear and Hummer, 2004). For example, in particle dynamics, population balance models are sometimes more effectively solved by the Monte Carlo method. Stochastic global optimization methods are very important algorithms for the solution of these types of problems. They have been successfully applied to solve challenging problems that cannot be solved using gradient based methods. Stochastic optimization methods have also been used in many algorithms, in which solution of optimization problems is part of the algorithm. Global stochastic optimization strategies have been utilized in learning phase of pattern recognition algorithms using fuzzy logic (Irizarry, 2005b) and neuro-fuzzy systems (Lin, 2008). These methods have been used for the optimization of complex engineering designs involving computational fluid mechanics such as aerodynamics applications (Duvigneau and Visonneau, 2004). Other applications include the determination of molecular structures, including protein structure prediction and protein-small molecule interactions among others (Sahinis, 2009). Batch scheduling problems are another type of problem were stochastic optimization can be very efficient (Liu et al., 2010). In particular, the solution of dynamic optimization problems is also of great industrial importance for process development and process optimization, since most processes are dynamic. In this type of problem an optimal profile function is sought (vs. an optimal value for a set of variables). For example, in a fed-batch fermenter, the feed-rate schedule is optimized to maximize production of antibiotics, vitamins, enzymes, and other products (Banga et al., 2003). Another example is the determination of optimal temperature profiles in crystallization processes to control crystal size distribution (Ma, Tafti and Braatz, 2002). Dynamic optimization is also of central importance to the application of process control
- Research Article
15
- 10.1007/s40998-020-00323-7
- Feb 7, 2020
- Iranian Journal of Science and Technology, Transactions of Electrical Engineering
Three degree of freedom (3 DOF) Hover Quad Copter (HQC) platforms are implemented for various missions in diverse scales from the micro to macro platforms. As HQC platforms scale down, micro platform requires rather robust and effective control techniques. This study investigates applicability of some stochastic optimization methods for tuning feedback gain control of HQC rotors and compares optimization results with results of linear quadratic regulator (LQR) method that has been widely used analytical method for optimal feedback gain control of HQCs. This study considers the utilization of two stochastic methods for tuning of HQCs. These methods are stochastic multi-parameter divergence optimization method (SMDO) and discrete stochastic optimization method (DSO). These methods are employed to optimize feedback gain coefficients of an experimental HQC test platform. Simulation and experimental results of SMDO and DSO methods are reported and compared with results of LQR method.
- Conference Article
20
- 10.5555/1218112.1218383
- Dec 3, 2006
This paper presents a stochastic traffic signal optimization method that consists of the CORSIM microscopic traffic simulation model and a heuristic optimizer. For the heuristic optimizer, the performance of three widely used optimization methods (i.e., genetic algorithm, simulated annealing and OptQuest Engine) was compared using a real world test corridor with 12 signalized intersections in Fairfax, Virginia, USA. The performance of the proposed stochastic optimization method was compared with an existing signal timing optimization program, SYNCHRO, under microscopic simulation environment. The results indicated that the genetic algorithm-based optimization method out-performs the SYNCHRO program as well as the other stochastic optimization methods in the optimization of traffic signal timings for the test corridor.
- Conference Article
27
- 10.1109/wsc.2006.322918
- Dec 1, 2006
This paper presents a stochastic traffic signal optimization method that consists of the CORSIM microscopic traffic simulation model and a heuristic optimizer. For the heuristic optimizer, the performance of three widely used optimization methods (i.e., genetic algorithm, simulated annealing and OptQuest Engine) was compared using a real world test corridor with 12 signalized intersections in Fairfax, Virginia, USA. The performance of the proposed stochastic optimization method was compared with an existing signal timing optimization program, SYNCHRO, under microscopic simulation environment. The results indicated that the genetic algorithm-based optimization method out-performs the SYNCHRO program as well as the other stochastic optimization methods in the optimization of traffic signal timings for the test corridor.
- Research Article
18
- 10.1038/s41598-021-90144-3
- May 21, 2021
- Scientific Reports
Deep learning applications require global optimization of non-convex objective functions, which have multiple local minima. The same problem is often found in physical simulations and may be resolved by the methods of Langevin dynamics with Simulated Annealing, which is a well-established approach for minimization of many-particle potentials. This analogy provides useful insights for non-convex stochastic optimization in machine learning. Here we find that integration of the discretized Langevin equation gives a coordinate updating rule equivalent to the famous Momentum optimization algorithm. As a main result, we show that a gradual decrease of the momentum coefficient from the initial value close to unity until zero is equivalent to application of Simulated Annealing or slow cooling, in physical terms. Making use of this novel approach, we propose CoolMomentum—a new stochastic optimization method. Applying Coolmomentum to optimization of Resnet-20 on Cifar-10 dataset and Efficientnet-B0 on Imagenet, we demonstrate that it is able to achieve high accuracies.
- Research Article
14
- 10.1016/j.fluid.2014.05.009
- May 15, 2014
- Fluid Phase Equilibria
A note on effective phase stability calculations using a Gradient-Based Cuckoo Search algorithm
- Research Article
9
- 10.3390/su14010553
- Jan 5, 2022
- Sustainability
Photovoltaic (PV) power generation has developed rapidly in recent years. Owing to its volatility and intermittency, PV power generation has an impact on the power quality and operation of the power system. To mitigate the impact caused by the PV generation, an energy storage (ES) system is applied to the PV plants. The capacity configuration and control strategy based on the stochastic optimization method have become an important research topic. However, the accuracy of the probability distribution model is insufficient and a stochastic optimization method is rarely used in a control strategy. In this paper, a stochastic optimization method for the energy storage system (ESS) configuration considering the self-regulation of the battery state of charge (SoC) is proposed. Firstly, to reduce the sampling error when typical scenarios of PV power are generated, a time-divided probability distribution model of the ultra-short-term predicted error of PV power is established. On this basis, to solve the problem that SoC reaches the threshold frequently, a self-regulation model of the SoC based on multiple scenarios is established, which can regulate the SoC according to rolling PV power prediction. A stochastic optimization configuration model of the energy storage system is constructed, which can reduce the impact of PV uncertainty on the configuration result. Finally, the proposed stochastic optimization method is validated. The fitting error of the time-divided probability distribution model is 15.61% lower than that of the t-distribution. The expected revenue of the optimal configuration in this paper is 8.86% higher than the scheme with a fixed probability distribution model, and 16.87% higher than without considering the stochastic optimization method.
- Research Article
7
- 10.11592/bit.120501
- Jan 1, 2012
- Biosystems and Information technology
Dynamic models give detailed information about the influence of many parameters on the behaviour of the biochemical process of interest. Parameter optimization of dynamic models is used in parameter estimation tasks and in design tasks. A drawback of the popular family of global stochastic optimization methods is the stochastic nature of the convergence of the best value of objective function to the global optimum or a value close to that. Therefore the optimization can take long time until a stable value of objective function is reached. Even then the risk of stagnation far from global optimum remains. That sets force to look for efficient approaches to reduce optimization time and discover cases of poor performance of optimization methods. Parallel optimization runs of identical optimization tasks can be used to reduce the impact of stochastic processes used in stochastic optimization methods. Consensus and stagnation criteria are proposed to terminate a set of parallel optimization runs when it is assessed that no significant improvements of the best value of the objective function are expected. Four automatically detectable cases of behaviour of a group of parallel optimization runs are analysed: 1) reaching of consensus criterion (consensus case), 2) stagnation of all optimization runs without reaching the consensus criterion (stagnation case), 3) stagnation at the initial value of the objective function, 4) lack of feasible solution. The proposed approach can be used automating the termination of optimization process when no further progress of the best value of objective function is expected. Suitability of particular optimization method with its settings for particular optimization task can be assessed analysing the dynamics of objective function's best values of parallel runs.
- Research Article
48
- 10.1109/tsp.2021.3092377
- Jan 1, 2021
- IEEE Transactions on Signal Processing
Stochastic compositional optimization generalizes classic (non-compositional) stochastic optimization to the minimization of compositions of functions. Each composition may introduce an additional expectation. The series of expectations may be nested. Stochastic compositional optimization is gaining popularity in applications such as reinforcement learning and meta learning. This paper presents a new Stochastically Corrected Stochastic Compositional gradient method (SCSC). SCSC runs in a single-time scale with a single loop, uses a fixed batch size, and guarantees to converge at the same rate as the stochastic gradient descent (SGD) method for non-compositional stochastic optimization. This is achieved by making a careful improvement to a popular stochastic compositional gradient method. It is easy to apply SGD-improvement techniques to accelerate SCSC. This helps SCSC achieve state-of-the-art performance for stochastic compositional optimization. In particular, we apply Adam to SCSC, and the exhibited rate of convergence matches that of the original Adam on non-compositional stochastic optimization. We test SCSC using the portfolio management and model-agnostic meta-learning tasks.
- Research Article
61
- 10.1016/j.scs.2022.103935
- Aug 1, 2022
- Sustainable Cities and Society
Optimal performance of hybrid energy system in the presence of electrical and heat storage systems under uncertainties using stochastic p-robust optimization technique
- Single Book
1
- 10.1115/1.862smo
- Apr 25, 2025
Stochastic Modeling and Optimization Methods for Critical Infrastructure Protection is a thorough exploration of mathematical models and tools that are designed to strengthen critical infrastructures against threats – both natural and adversarial. Divided into two volumes, this first volume examines stochastic modeling across key economic sectors and their interconnections, while the second volume focuses on advanced mathematical methods for enhancing infrastructure protection. The book covers a range of themes, including risk assessment techniques that account for systemic interdependencies within modern technospheres, the dynamics of uncertainty, instability and system vulnerabilities. The book also presents other topics such as cryptographic information protection and Shannon’s theory of secret systems, alongside solutions arising from optimization, game theory and machine learning approaches. Featuring research from international collaborations, this book covers both theory and applications, offering vital insights for advanced risk management curricula. It is intended not only for researchers, but also educators and professionals in infrastructure protection and stochastic optimization.
- Research Article
4
- 10.3390/en18195154
- Sep 28, 2025
- Energies
The paper investigates the design and operation of microgrid arrangements, with a focus on renewable power systems, system architectures, and storage solutions. The research evaluates stochastic and multi-objective optimization methods to show how demand response systems improve operational flexibility. The study evaluates 183 journal articles to select those that address microgrid design in conjunction with optimization models and demand response approaches. The articles are classified into three essential categories, which include microgrid design optimization methods and demand response integration. The review establishes that microgrid performance depends on three fundamental design parameters, which include energy generation systems, storage capabilities, and load demand control mechanisms. The review demonstrates that advanced optimization approaches, such as stochastic and multi-objective optimization methods, offer effective solutions for managing renewable energy variability. The paper demonstrates that demand response strategies are crucial for reducing costs and enhancing system flexibility. However, current published research falls short of establishing an integrated system that combines real-time demand response with stochastic optimization. This integration, while not yet fully realized, is suggested as a critical advancement for ensuring both system performance optimization and long-term sustainability. Therefore, this paper calls for further research to develop resilient hybrid renewable microgrids that integrate flexibility with sustainability through advanced optimization models and demand response strategies.
- Single Book
1
- 10.1115/1.862sma
- Apr 25, 2025
Stochastic Modeling and Optimization Methods for Critical Infrastructure Protection is a thorough exploration of mathematical models and tools that are designed to strengthen critical infrastructures against threats both natural and adversarial. Divided into two volumes, this first volume examines stochastic modeling across key economic sectors and their interconnections, while the second volume focuses on advanced mathematical methods for enhancing infrastructure protection. The book covers a range of themes, including risk assessment techniques that account for systemic interdependencies within modern technospheres, the dynamics of uncertainty, instability and system vulnerabilities. The book also presents other topics such as cryptographic information protection and Shannon's theory of secret systems, alongside solutions arising from optimization, game theory and machine learning approaches. Featuring research from international collaborations, this book covers both theory and applications, offering vital insights for advanced risk management curricula. It is intended not only for researchers, but also educators and professionals in infrastructure protection and stochastic optimization.
- Book Chapter
29
- 10.1016/b978-008044046-0.50582-0
- Jan 1, 2003
- Computational Fluid and Solid Mechanics 2003
Solution of implicitly formulated inverse heat transfer problems with hybrid methods
- Research Article
43
- 10.1016/j.cma.2007.07.003
- Jul 25, 2007
- Computer Methods in Applied Mechanics and Engineering
Stochastic design optimization: Application to reacting flows