DeepAR: Probabilistic forecasting with autoregressive recurrent networks
Probabilistic forecasting, i.e., estimating a time series’ future probability distribution given its past, is a key enabler for optimizing business processes. In retail businesses, for example, probabilistic demand forecasts are crucial for having the right inventory available at the right time and in the right place. This paper proposes DeepAR, a methodology for producing accurate probabilistic forecasts, based on training an autoregressive recurrent neural network model on a large number of related time series. We demonstrate how the application of deep learning techniques to forecasting can overcome many of the challenges that are faced by widely-used classical approaches to the problem. By means of extensive empirical evaluations on several real-world forecasting datasets, we show that our methodology produces more accurate forecasts than other state-of-the-art methods, while requiring minimal manual work.
- Research Article
707
- 10.1016/0016-0032(78)90124-2
- Jan 1, 1978
- Journal of the Franklin Institute
Optimum systems control: by A. P. Sage and C. C. White, III. 413 pages, illustrations, [formula omitted] in. Prentice-Hall, Englewood Cliffs, N.J., 1977. Price $22.50 (approx. £13.00)
- Research Article
11
- 10.1109/tnnls.2022.3190984
- Apr 1, 2024
- IEEE Transactions on Neural Networks and Learning Systems
We introduce a novel method to estimate the causal effects of an intervention over multiple treated units by combining the techniques of probabilistic forecasting with global forecasting methods using deep learning (DL) models. Considering the counterfactual and synthetic approach for policy evaluation, we recast the causal effect estimation problem as a counterfactual prediction outcome of the treated units in the absence of the treatment. Nevertheless, in contrast to estimating only the counterfactual time series outcome, our work differs from conventional methods by proposing to estimate the counterfactual time series probability distribution based on the past preintervention set of treated and untreated time series. We rely on time series properties and forecasting methods, with shared parameters, applied to stacked univariate time series for causal identification. This article presents DeepProbCP, a framework for producing accurate quantile probabilistic forecasts for the counterfactual outcome, based on training a global autoregressive recurrent neural network model with conditional quantile functions on a large set of related time series. The output of the proposed method is the counterfactual outcome as the spline-based representation of the counterfactual distribution. We demonstrate how this probabilistic methodology added to the global DL technique to forecast the counterfactual trend and distribution outcomes overcomes many challenges faced by the baseline approaches to the policy evaluation problem. Oftentimes, some target interventions affect only the tails or the variance of the treated units' distribution rather than the mean or median, which is usual for skewed or heavy-tailed distributions. Under this scenario, the classical causal effect models based on counterfactual predictions are not capable of accurately capturing or even seeing policy effects. By means of empirical evaluations of synthetic and real-world datasets, we show that our framework delivers more accurate forecasts than the state-of-the-art models, depicting, in which quantiles, the intervention most affected the treated units, unlike the conventional counterfactual inference methods based on nonprobabilistic approaches.
- Research Article
5
- 10.1080/24751839.2023.2250113
- Sep 6, 2023
- Journal of Information and Telecommunication
Ethereum is a major public blockchain. Besides being the second-largest digital currency by market capitalization for its cryptocurrency, the Ether (Ξ), it is also the foundation of Web3 and decentralized applications, or DApps, that are fuelled by Smart Contracts. At the time of this writing, Ethereum still uses Proof of Work (PoW) consensus algorithm to ensure the integrity of the blockchain and to prevent double spend. PoW requires the participation of miners, who are incentivized to assemble blocks of transactions by being rewarded with cryptocurrency paid by transaction originators and by the blockchain network itself via newly minted Ξ. Network fees for transaction submissions are called gas, by analogy to the fuel used by cars, and are negotiable. They are also highly volatile and hence it is critical to predict the direction they are heading into, so that one can time transaction submissions, when feasible. There have been several efforts to predict gas prices, including usage of large Mempools, analysis of committed blocks, and more recent ones using Facebook's Prophet model [Taylor, S. J., & Letham, B. (2017). Forecasting at scale. PeerJ Preprints, 5, e3190v2. https://doi.org/10.7287/peerj.preprints.3190v2]. In this study, we introduce an innovative approach that employs the DeepAR [Salinas, D., Flunkert, V., Gasthaus, J., & Januschowski, T. (2020). Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3), 1181–1191. https://doi.org/10.1016/j.ijforecast.2019.07.001] model, known for its superior forecasting accuracy over conventional methods by virtue of its ability to learn from multiple related time series. This methodology not only offers immediate advantages but also holds promise for ongoing enhancements. We substantiate our claims through empirical testing, utilizing data extracts from the Ethereum blockchain and cryptocurrency price feeds. This document is an extended version of our ICCS 2022 paper on the same topic. In this paper, we dive deeper into the internals of DeepAR forecasting algorithm [Salinas, D., Flunkert, V., Gasthaus, J., & Januschowski, T. (2020). Deepar: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting, 36(3), 1181–1191. https://doi.org/10.1016/j.ijforecast.2019.07.001], analyse the correlation between the on-chain/off-chain sample data, and describe additional experiments that empirically prove our findings and, finally, perform a comparison of our outputs with those from the Prophet [Taylor, S. J., & Letham, B. (2017). Forecasting at scale. PeerJ Preprints, 5, e3190v2. https://doi.org/10.7287/peerj.preprints.3190v2] model.
- Research Article
- 10.5455/ajvs.265570
- Jan 1, 2025
- Alexandria Journal of Veterinary Sciences
Monitoring animal productivity is a crucial aspect of farm management and decision-making. Sheep breeding is of utmost importance to the agricultural and veterinary industries. However, unlike veterinary epidemiology, applying forecasting models to predict animal productivity remains less explored, and developing reliable predictive modelling is one of the necessary strategies in this regard. This study addresses this gap by focusing on lamb production. To forecast lamb production trends accurately, this study examines the effectiveness of traditional time-series models ARIMA and ETS alongside the neural network autoregressive (NNAR) model, a machine learning-based approach, to identify the most suitable forecasting method. Also, the study aimed to address a critical question about the reliability of predicting productivity with short-time series. The models are trained and evaluated using yearly lamb production data collected from 2003 to 2022. The model exhibiting the lowest corrected Akaike Information Criterion (AICc) was identified as the optimal forecasting model. Its accuracy was confirmed using the lowest mean absolute percentage error (MAPE) and root mean square error (RMSE). Our findings indicate that among the evaluated models, NNAR yielded the most accurate forecasts, reflected in its minimal MAPE and RMSE values compared to ARIMA and ETS. This highlights its superior predictive performance for short time series in annual sheep production. While ARIMA outperformed ETS on most evaluation metrics, NNAR consistently delivered the best results. Notably, NNAR maintained high accuracy despite the small sample size (20 observations), whereas ARIMA may require a larger dataset to perform optimally. Overall, these models are suitable for time series analysis and hold promise for broader application in animal science to enhance forecasting and decision-making.
- Research Article
- 10.3897/aca.8.e151700
- May 28, 2025
- ARPHA Conference Abstracts
In recent years, alternating drought and extreme precipitation events have highlighted the need for subseasonal to seasonal forecasts of the terrestrial water cycle. In particular, predictions of the impacts of dry and wet extremes on groundwater resources are crucial to assess the impacts of groundwater scarcity and excess on ecosystem dynamics as well as to provide stakeholders in agriculture, forestry, the water sector, and other fields with information supporting the sustainable use of these resources. In this context, we calculate four times per year seasonal probabilistic hydrological forecasts at 0.6 km resolution, from the surface down to 60 m depth, for the upcoming seven months over Germany and surrounding regions (hydrological Germany). These forecasts are generated using the integrated, physics-based hydrological model ParFlow/CLM (Kuffour et al. 2020) with the setup described in Belleflamme et al. (2023), forced by 50 ensemble members of the SEAS5 seasonal forecast from the European Centre for Medium-Range Weather Forecasts (ECMWF). The predicted evolution of the terrestrial water cycle is released at the beginning of each meteorological season for the total subsurface water storage anomaly in our experimental Water Resources Bulletin (https://adapter-projekt.de/bulletin/index.html). To evaluate our forecasts, we evaluated six 7-months probabilistic forecasts covering the vegetation period (March to September) for the years 2018 to 2023 with a reference long-term historical time series based on the same ParFlow/CLM setup. The forecast skill was assessed by comparing these seasonal forecasts to a climatology-based 10-member pseudo-forecast over the 2013–2023 period (using the leave-one-out method), extracted from the reference time series. The monthly Continuous Ranked Probability Skill Score (CRPSS), which evaluates the ensemble distribution based on daily water table depth data, indicates that the probabilistic forecast outperforms the climatology-based pseudo-forecast in most regions, except in 2018 and, to a lesser extent, in 2020 and 2022 (see Fig. 1). This can be attributed to an under-representation of extremely dry members in the ensemble, combined with the memory effect of the initial conditions at increasing soil depths. For example, while March 2018 started with a slightly below-average water table depth and experienced a strong meteorological drought leading to an agricultural drought and eventually an increase in water table depth (i.e., groundwater depletion), the initial water table depth anomaly in March 2019 was already positive, with a less pronounced precipitation deficit during the vegetation period. This resulted in a much higher forecast skill, because of the memory effect accurately simulated with the physics-based model. Notably, the forecast skill only slightly decreases with increasing lead time, both for precipitation and water table depth. The analysis of the Relative Operating Characteristic Skill Score (ROCSS) for the upper quintile of the water table depth distribution assesses whether positive water table depth anomalies (i.e., droughts) are adequately represented within the probabilistic forecast ensemble. The results are consistent with those of the CRPSS, showing lower skill in 2018. Nevertheless, the ROCSS analysis overall indicates high skill for the probabilistic forecast, while the climatology-based pseudo-forecast demonstrates no skill. This means that the forecast ensemble covers a sufficiently wide range of possible scenarios to include extreme events of the order of magnitude as the droughts experienced in central Europe over recent years. To conclude, this evaluation confirms that the dry conditions experienced in central Europe in recent years were captured within the probabilistic forecast, underlining the added value of these forecasts and their potential usefulness for predicting and assessing the impact of groundwater fluctuations on ecosystem functioning and dynamics.
- Research Article
49
- 10.3390/s21010014
- Dec 22, 2020
- Sensors
With increased urbanization, accidents related to slope instability are frequently encountered in construction sites. The deformation and failure mechanism of a landslide is a complex dynamic process, which seriously threatens people’s lives and property. Currently, prediction and early warning of a landslide can be effectively performed by using Internet of Things (IoT) technology to monitor the landslide deformation in real time and an artificial intelligence algorithm to predict the deformation trend. However, if a slope failure occurs during the construction period, the builders and decision-makers find it challenging to effectively apply IoT technology to monitor the emergency and assist in proposing treatment measures. Moreover, for projects during operation (e.g., a motorway in a mountainous area), no recognized artificial intelligence algorithm exists that can forecast the deformation of steep slopes using the huge data obtained from monitoring devices. In this context, this paper introduces a real-time wireless monitoring system with multiple sensors for retrieving high-frequency overall data that can describe the deformation feature of steep slopes. The system was installed in the Qili connecting line of a motorway in Zhejiang Province, China, to provide a technical support for the design and implementation of safety solutions for the steep slopes. Most of the devices were retained to monitor the slopes even after construction. The machine learning Probabilistic Forecasting with Autoregressive Recurrent Networks (DeepAR) model based on time series and probabilistic forecasting was introduced into the project to predict the slope displacement. The predictive accuracy of the DeepAR model was verified by the mean absolute error, the root mean square error and the goodness of fit. This study demonstrates that the presented monitoring system and the introduced predictive model had good safety control ability during construction and good prediction accuracy during operation. The proposed approach will be helpful to assess the safety of excavated slopes before constructing new infrastructures.
- Research Article
22
- 10.1016/j.energy.2024.131966
- Jun 11, 2024
- Energy
Hybrid model based on similar power extraction and improved temporal convolutional network for probabilistic wind power forecasting
- Research Article
98
- 10.1016/j.solener.2018.06.100
- Jul 4, 2018
- Solar Energy
Day-ahead probabilistic PV generation forecast for buildings energy management systems
- Conference Article
2
- 10.1109/isgt.2019.8791617
- Feb 1, 2019
Photovoltaic (PV) generation forecasting is critical in advancing the PV penetration level in power systems. However, the inherent uncertainties of PV generation can hardly be quantified by traditional point forecast methods as point forecasting methods can only provide one predicted value at each time step. Due to its superior ability of measuring uncertainties, probabilistic forecasting, which can be in the form of quantile forecasts, is drawing increasing attention. Nevertheless, due to the extensive stochastic nature of PV generation, the existing probabilistic forecasting methods are not sufficiently reliable to quantify the uncertainties of PV generation. This paper adapts quantile determination (QD) to improve the reliability of probabilistic PV generation forecasting. Based on the modified QD, a new method for probabilistic PV generation forecasting is proposed. A one-hour-ahead PV generation forecasting case study reveals that with the proposed QD, the reliability of probabilistic PV forecasting is significantly improved. Moreover, it is also shown that the proposed probabilistic PV generation forecasting method can conduct much more accurate and reliable quantile forecasts, compared with existing methods.
- Conference Article
7
- 10.2118/214769-ms
- Oct 9, 2023
The majority of production forecasting methods currently used are point forecasting methods developed in the setting of individual well forecasting. For an actual oilfield, instead of needing to predict individual production time series, one is faced with forecasting thousands of related time series and the uncertainty can be assessed. The objective of this work is to enable global modeling and probabilistic forecasting of a large number of related production time series using Deep Autoregressive Recurrent Neural Networks (DeepAR). The DeepAR model consists of three parts. First, the auxiliary data such as static classification covariates and dynamic covariates are encoded. Second, establish a forward model based on an autoregressive recurrent neural network. Third, the normal distribution is defined as the output distribution function. And the variance and mean are obtained by solving the maximum log-likelihood function using the gradient descent algorithm. We demonstrate how the application of DeepAR to forecasting can overcome many of the challenges(e.g. frequent well shut-in and opening, probabilistic prediction, classification prediction) that are faced by widely-used classical approaches to the problem. In this work, history fitting and prediction were performed on a dataset from more than 2000 tight gas reservoir wells in the Ordos Basin, China. The DeepAR and conventional methods were tested and compared based on the datasets. We show through extensive empirical evaluation on several real-world forecasting data sets accuracy improvements of around 30% compared to RNN-based networks. In the case of frequent well shut-ins and openings, the RNN-based network structure cannot capture the fast pressure response and extreme fluctuations, which eventually leads to high errors. In contrast, DeepAR is more stable to frequent or significant well variations, can learn different dynamic and static category features, generates calibrated probabilistic forecasts with high accuracy, and can learn complex patterns such as seasonality and uncertainty growth over time from the data. This study provides more general production forecasting and analysis of production dynamics methods from a big data perspective. Instead of performing costly well tests or shut-ins, reservoir engineers can extract valuable long-term reservoir performance information from predictions estimated by DeepAR trained on an extensive collection of related production time series data.
- Research Article
26
- 10.1016/j.jhydrol.2023.130498
- Nov 21, 2023
- Journal of Hydrology
Generative deep learning for probabilistic streamflow forecasting: Conditional variational auto-encoder
- Research Article
28
- 10.1287/msom.2023.1210
- Apr 11, 2023
- Manufacturing & Service Operations Management
Problem definition: We study the estimation of the probability distribution of individual patient waiting times in an emergency department (ED). Whereas it is known that waiting-time estimates can help improve patients’ overall satisfaction and prevent abandonment, existing methods focus on point forecasts, thereby completely ignoring the underlying uncertainty. Communicating only a point forecast to patients can be uninformative and potentially misleading. Methodology/results: We use the machine learning approach of quantile regression forest to produce probabilistic forecasts. Using a large patient-level data set, we extract the following categories of predictor variables: (1) calendar effects, (2) demographics, (3) staff count, (4) ED workload resulting from patient volumes, and (5) the severity of the patient condition. Our feature-rich modeling allows for dynamic updating and refinement of waiting-time estimates as patient- and ED-specific information (e.g., patient condition, ED congestion levels) is revealed during the waiting process. The proposed approach generates more accurate probabilistic and point forecasts when compared with methods proposed in the literature for modeling waiting times and rolling average benchmarks typically used in practice. Managerial implications: By providing personalized probabilistic forecasts, our approach gives low-acuity patients and first responders a more comprehensive picture of the possible waiting trajectory and provides more reliable inputs to inform prescriptive modeling of ED operations. We demonstrate that publishing probabilistic waiting-time estimates can inform patients and ambulance staff in selecting an ED from a network of EDs, which can lead to a more uniform spread of patient load across the network. Aspects relating to communicating forecast uncertainty to patients and implementing this methodology in practice are also discussed. For emergency healthcare service providers, probabilistic waiting-time estimates could assist in ambulance routing, staff allocation, and managing patient flow, which could facilitate efficient operations and cost savings and aid in better patient care and outcomes.Supplemental Material: The online supplement is available at https://doi.org/10.1287/msom.2023.1210 .
- Research Article
293
- 10.1016/j.ijepes.2021.107818
- Dec 9, 2021
- International Journal of Electrical Power & Energy Systems
Short-term load forecasting based on LSTM networks considering attention mechanism
- Research Article
- 10.1038/s41598-026-53844-2
- Jun 3, 2026
- Scientific reports
Accurate and reliable forecasting of wind power generation is a cornerstone for ensuring the operational flexibility, economic efficiency, and grid stability of modern power systems with high renewable penetration. While recent advances in deep learning have improved point and probabilistic forecasting, most existing methods still emphasize error minimization at isolated horizons or locations, neglecting the spatiotemporal dependencies, uncertainty calibration, and temporal coherence that are critical in real-world decision-making contexts. In this study, we propose a novel adaptive spatiotemporal graph neural network (ST-GNN) framework for multi-horizon probabilistic wind power forecasting, designed to jointly optimize deterministic accuracy, distributional sharpness, and temporal stability. The framework integrates site-specific meteorological inputs and SCADA data into a dynamically evolving graph structure, where edges are adaptively reweighted to capture changing inter-site correlations driven by wind climatology. A unified multi-objective loss function is formulated, combining mean error minimization with probabilistic metrics such as the Continuous Ranked Probability Score (CRPS) and a temporal smoothness regularizer that mitigates spurious fluctuations across consecutive horizons. Comprehensive case studies are conducted using two years of high-resolution data from a wind farm cluster consisting of twelve geographically diverse sites. Results demonstrate that the proposed model achieves up to 13% reduction in RMSE at 1-4 hour horizons and maintains improvements above 5% at 24-hour horizons when compared with state-of-the-art baselines. Probabilistic evaluations indicate a reduction in calibration error, with empirical coverage deviations generally within 3% across central quantiles and uncertainty estimates that adapt to different meteorological regimes. Moreover, the temporal rank-order consistency of predictions remains significantly higher than existing models, reducing dispatch-related instabilities and improving multi-site coordination. These findings suggest that the integration of spatiotemporal graph structures with probabilistic training objectives can provide a flexible and potentially generalizable paradigm.
- Conference Article
6
- 10.1109/pmaps.2016.7764081
- Oct 1, 2016
Model selection is an important step for both point and probabilistic load forecasting. In the point load forecasting literature and practices, point error measures, such as mean absolute percentage error (MAPE), are often used for model selection. On the other hand, many probabilistic load forecasting methodologies rely on the model selection mechanism developed for point load forecasting. In other words, the models for probabilistic load forecasting are selected to minimize point error measures rather than probabilistic ones, such as quantile score. Intuitively, selecting models for probabilistic forecasting based on a point error measure is less computationally intensive and less accurate than its counterpart. The practical question is whether we can gain significant accuracy by taking the more computationally intensive route. This paper presents a comparative study on model selection for probabilistic load forecasting, using point and probabilistic error measures respectively. The data for the case study is from the load forecasting track of the Global Energy Forecasting Competition 2014. We find that the two model selection mechanisms indeed return different underlying models. While on average, the models from quantile score based model selection method can lead to more accurate probabilistic forecasts, the improvement over the MAPE based model selection method is marginal.