Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Accuracy of rainfall prediction using deep learning based on a recurrent neural network with an LSTM layer method

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Accuracy of rainfall prediction using deep learning based on a recurrent neural network with an LSTM layer method

Similar Papers
  • Research Article
  • 10.47065/bulletincsr.v5i5.735
Optimalisasi Akurasi Prediksi Curah Hujan Bulanan Menggunakan Deep Learning
  • Aug 26, 2025
  • Bulletin of Computer Science Research
  • Muhammad Ikrom Yafik + 1 more

The Province of Lampung exhibits high rainfall variability influenced by various atmospheric dynamics such as the Asian Monsoon, Australian Monsoon, El Niño–Southern Oscillation (ENSO), and the Indian Ocean Dipole (IOD). Accurate rainfall prediction is crucial across multiple sectors, including agriculture, water resource management, and hydrometeorological disaster mitigation. However, prediction methods commonly used in the region are still dominated by statistical approaches or conventional machine learning techniques, which often struggle to capture long-term temporal patterns in rainfall data. On the other hand, deep learning technologies such as the Recurrent Neural Network (RNN) and Gated Recurrent Unit (GRU) offer better capabilities in modeling time series data, yet no specific comparative evaluation has been conducted for rainfall prediction in the Lampung Province. Comparing these two methods is important because the architectural characteristics of RNN and GRU differ in handling long-term dependencies, and selecting the right model can directly impact prediction accuracy and the effectiveness of decision-making in affected sectors. This study aims to implement and compare the performance of RNN and GRU in predicting monthly rainfall in Lampung Province using data from 80 rain gauges distributed across 15 districts/cities over the period from January 1991 to February 2025. The results show that the RNN model outperforms the GRU model, with lower RMSE (115.61 vs. 119.50), smaller MAE (86.94 vs. 91.28), and higher R² (0.35 vs. 0.30). Predictions for the period from March 2025 to February 2026 reveal a clear seasonal pattern, with minimum rainfall occurring in August 2025 (peak dry season) and maximum rainfall in January 2026 (peak rainy season). This study demonstrates that RNN is more effective than GRU in capturing the temporal patterns of rainfall, making it more recommended for long-term prediction applications.

  • Research Article
  • 10.1504/ijw.2025.10073347
Accuracy of rainfall prediction using deep learning based on a recurrent neural network with an LSTM layer method
  • Jan 1, 2025
  • International Journal of Water
  • Rindra Yusianto + 3 more

Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas.

  • Research Article
  • 10.46632/cset/3/4/3
A Comparative Study of Recurrent Neural Network (RNN) with Gray Relational Analysis for Temporal Data
  • Dec 6, 2025
  • Computer Science, Engineering and Technology

A Recurrent Neural Network (RNN) is a specialized form of neural network that is adept at handling sequential data by retaining information from prior inputs. In contrast to conventional feedforward neural networks, RNNs incorporate loops in their architecture, allowing them to leverage data from previous time steps to affect the current output. This characteristic renders RNNs especially effective for applications that involve sequences, including time-series forecasting, natural language processing, and speech recognition. A fundamental component of RNNs is their hidden state, which acts as a dynamic memory that is refreshed with each incoming input. This allows RNNs to capture dependencies across time steps, which is crucial for understanding context in sequences. In language modeling, the interpretation of a word often relies on the words that come before it, a task that Recurrent Neural Networks (RNNs) handle well. However, RNNs struggle with issues like vanishing gradients, which hinder their ability to capture long-range dependencies. To overcome this, models such as Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs) were introduced. These models incorporate gates that regulate the flow of information, allowing them to better learn long-term dependencies. RNNs remain a powerful tool for working with sequential data, facilitating the modeling of temporal relationships, but their effectiveness depends on careful design and optimization. Research significance: Recurrent Neural Networks (RNNs) hold significant research value because of their capacity to simulate temporal and sequential data, which is essential in many fields. They are frequently employed in natural language processing for tasks such as sentiment analysis, language translation, and text generation. In time-series analysis, RNNs enable accurate forecasting in finance, healthcare, and climate modeling. They also are essential in speech recognition and video processing, handling dependencies across time steps. Research focuses on improving RNNs, addressing challenges like vanishing gradients, and enhancing efficiency through architectures like LSTMs and GRUs, solidifying their relevance in advancing AI and machine learning applications. Methodology: A technique for analyzing the relationships between several variables, particularly in situations when data is limited or unclear, is called gray relational analysis, or GRA. In order to comprehend the relationships between variables, it evaluates how similar or different they are. GRA aids decision-makers in identifying critical factors, prioritizing actions, and improving processes in complex fields like engineering, finance, and management. By converting both qualitative and quantitative data into gray numbers, GRA addresses uncertainty and provides valuable insights for problem-solving, decision-making, and performance improvement, leading to more informed and effective strategies. Alternative taken as Simple RNN, LSTM, GRU, Bidirectional RNN, Deep RNN, Vanilla RNN, Echo State Network, Attention-based RNN, Transformer RNN, GRU with Attention. Evaluation preference taken as Prediction Accuracy, Model Robstness, Learning Efficiency, Training Time, Complexity. Attention-based RNN has the lowest score, Deep RNN has the highest rank, according to the results.

  • Research Article
  • 10.1175/jhm-d-24-0015.1
Virtual Orbital Simulation and OSSE of Radiometer-Equipped Small Satellite Constellation for River-Basin-Scale Rainfall Prediction
  • May 1, 2025
  • Journal of Hydrometeorology
  • Rie Seto + 1 more

Small satellite constellations (SSCs) equipped with microwave radiometers are rapidly advancing for Earth observation. The application of such technology to hydrometeorology is imminent. This study proposes a method to evaluate the impact of measurements from radiometer-equipped SSCs on river-basin-scale rainfall prediction. By combining virtual orbital simulation and observing system simulation experiments (OSSE), the method assimilates virtual but realistic measurements. It also examines the relationship between SSC configurations and prediction accuracy, thus providing the SSC design community with information for optimizing SSC setups for hydrometeorological applications. Results show that frequent assimilation of SSC data improves river-basin-scale precipitation forecasts by increasing opportunities to generate accurate initial conditions. Substantial improvements were observed also in water vapor and temperature fields, with assimilation effects lasting for over 6 h. However, increasing the number of satellites does not always lead to better prediction accuracy. Furthermore, the study reveals that evenly spaced observation intervals lead to more stable and accurate rainfall predictions, highlighting the importance of both frequency and timing in SSC configuration. These findings emphasize the need for careful consideration of SSC design to maximize the benefits of this technology for rainfall prediction. Significance Statement This study introduces a new method to validate river-basin-scale rainfall predictions by assimilating virtual measurements from a radiometer-equipped small satellite constellation (SSC) combining virtual orbital simulations and observing system simulation experiment (OSSE). It also proposes a framework for recommending optimal SSC configurations to achieve desirable prediction accuracy. The results show that frequent assimilation of SSC data has great potential to enhance rainfall prediction accuracy. The results also highlight the importance of carefully designing SSC configurations to maximize their effectiveness in improving rainfall forecasts.

  • PDF Download Icon
  • Preprint Article
  • Cite Count Icon 1
  • 10.21203/rs.3.rs-3375438/v1
Application of Maximum Overlap Discrete Wavelet Transform and Machine Learning to Improved Daily Rainfall Prediction Modeling
  • Sep 27, 2023
  • Research Square
  • Kübra Küllahci + 1 more

Rainfall is an important phenomenon for various aspects of human life and the environment. Accurate prediction of rainfall is crucial for a wide range of sectors, including agriculture, water resources management, energy production, disaster management, and many more. The ability to predict rainfall in an accurate fashion enables stakeholders to make informed decisions and take necessary actions to mitigate the impacts of natural disasters, water scarcity, and other issues related to rainfall. In addition, advances in rainfall prediction technologies have the potential to contribute to sustainable water management and the preservation of water resources by providing the necessary information for decision-makers to plan and implement effective water management strategies. Hence, it is important to continuously improve the accuracy of rainfall prediction. In this paper, the integration of the Maximum Overlap Discrete Wavelet Transform (MODWT) and machine learning algorithms for daily rainfall prediction is proposed. The main objective of this study is to investigate the potential of combining MODWT with various machine-learning algorithms to increase the accuracy of rainfall prediction and extend the forecast time horizon to three days. In addition, the performances of the proposed hybrid models are contrasted with the models hybridized with commonly used discrete wavelet transform (DWT) algorithms in the literature. For this, daily rainfall raw data from 3 rainfall observation stations located in Türkiye are used. The results show that the proposed hybrid MODWT models can effectively improve the accuracy of precipitation forecasting, based on model evaluation measures such as mean square error (MSE) and Nash-Sutcliffe coefficient of efficiency (CE). Accordingly, it can be concluded that the integration of MODWT and machine learning algorithms have the potential to revolutionize the field of daily rainfall prediction.

  • Research Article
  • Cite Count Icon 9
  • 10.1002/joc.8530
Maximizing daily rainfall prediction accuracy with maximum overlap discrete wavelet transform‐based machine learning models
  • Jun 12, 2024
  • International Journal of Climatology
  • Kübra Küllahcı + 1 more

Rainfall is an important phenomenon for various aspects of human life and the environment. Accurate prediction of rainfall is crucial for a wide range of sectors, including agriculture, water resources management, energy production, disaster management and many more. The ability to predict rainfall in an accurate fashion enables stakeholders to make informed decisions and take necessary actions to mitigate the impacts of natural disasters, water scarcity and other issues related to rainfall. In addition, advances in rainfall prediction technologies have the potential to contribute to sustainable water management and the preservation of water resources by providing the necessary information for decision‐makers to plan and implement effective water management strategies. Hence, it is important to continuously improve the accuracy of rainfall prediction. In this paper, the integration of the maximum overlap discrete wavelet transform (MODWT) and machine learning algorithms for daily rainfall prediction is proposed. The main objective of this study is to investigate the potential of combining MODWT with various machine‐learning algorithms to increase the accuracy of rainfall prediction and extend the forecast time horizon to 3 days. In addition, the performances of the proposed hybrid models are contrasted with the models hybridized with commonly used discrete wavelet transform (DWT) algorithms in the literature. For this, daily rainfall raw data from three rainfall observation stations located in Turkey are used. The results show that the proposed hybrid MODWT models can effectively improve the accuracy of precipitation forecasting, based on model evaluation measures such as mean square error (MSE) and Nash‐Sutcliffe coefficient of efficiency (CE). Accordingly, it can be concluded that the integration of MODWT and machine learning algorithms have the potential to revolutionize the field of daily rainfall prediction.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/piers53385.2021.9695104
Analyzing Impact of Time on Early Detection of Rainfall Event
  • Nov 21, 2021
  • Muhammad Salman Pathan + 4 more

Rainfall is a critical feature of a climatic system, which has a chaotic impact on agriculture, water resource management and biological systems. An early and accurate prediction of rainfall is a very important task and has vital effects on human life. However rainfall prediction is a challenging task in meteorology. Rainfall data mostly have high inconstancy and irregular patterns which are rare in other time series data. The rainfall data changes greatly with time, so the time factor has a high importance in such time series data. In order to develop efficient forecast models, one should deeply analyze the effect of time on prediction accuracy. Therefore, in this paper, we analyze how early we can predict a rainfall event with accuracy. We have performed the experiments on a 5-year daily rainfall data obtained from National Oceanic and Atmospheric Administration (NOAA) at 1-, 2-, 3-, and 4-day prior forecasting horizons using a machine learning technique to observe the trends in the prediction accuracy. Furthermore, we have also identified some most important input features from the dataset which plays a major role in the prediction of a rainfall event. Our results conclude that average wind speed and minimum temperature are most important weather variables in classifying rainfall events. We also observe that the forecasting error gradually increases with increasing lead times.

  • Conference Article
  • Cite Count Icon 30
  • 10.1109/icirca.2018.8597421
Prediction of Rainfall Using Artificial Neural Network
  • Jul 1, 2018
  • A Kala + 1 more

Rainfall prediction is an important and challenging task in meteorology. Rainfall is predicted using different models with their combination, observation, trends of knowledge and patterns. Rainfall can be predicted using various machine learning techniques. In this paper, Artificial Neural Network (ANN) such as Feed Forward Neural Network (FFNN) model is built for predicting the rainfall. Artificial neural networks (ANN) are the valuable and attractive soft computing method for prediction. ANN is based on self-adaptive mechanism in which the model learns from historical data capture functional relationships between data and make predictions on current data. The accurate prediction of rainfall is a major criterion for managing the water resources. The prediction accuracy is measured using confusion matrix and RMSE. The results show that the prediction model based on ANN indicates acceptable accuracy.

  • Book Chapter
  • 10.4018/979-8-3373-2387-9.ch006
Accurate Rainfall Prediction With Integrated BI LSTM in Deep Learning
  • Jul 11, 2025
  • Usharani Bhimavarapu

Accurate rainfall prediction is essential for agricultural planning and water resource management, particularly in regions like Andhra Pradesh, where the economy relies heavily on monsoon-dependent agriculture. Traditional statistical and empirical models, such as linear regression and ARIMA, have been commonly employed for rainfall forecasting in the region. However, these models often fail to account for the complex, non-linear relationships between variables that influence rainfall, resulting in inaccurate predictions, especially during extreme weather events. This study proposes a hybrid deep learning model combining Bi-LSTM (Bidirectional Long Short-Term Memory) and GRU (Gated Recurrent Unit) to improve rainfall prediction accuracy. The model leverages the strengths of both Bi-LSTM and GRU networks to capture both short- and long-term dependencies in historical rainfall data.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/yac53711.2021.9486432
TCM Medical Record Analysis Algorithm Based on Recurrent Convolutional Neural Network
  • May 28, 2021
  • Yu Zhang + 4 more

The most important step in the process of medical record analysis in TCM is the classification of medical records. The biggest challenge of medical record classification is to perceive the correlation between context words and find keywords, and make judgments based on the keyword information. In this article, we propose a TCM medical record analysis algorithm based on recurrent convolutional neural network, which introduces a maximum pooling layer in the recurrent neural network, and uses it to determine the words that play an important role in text classification to capture the key components of the text. Experimental results show that recurrent convolutional neural network achieves better results than attention recurrent neural network and traditional recurrent neural network. In addition, recurrent convolutional neural network is more than twice as fast as them in terms of training speed.

  • Peer Review Report
  • 10.7554/elife.83035.sa0
Editor's evaluation: Neural population dynamics of computing with synaptic modulations
  • Jan 8, 2023
  • Gianluigi Mongillo

Article Figures and data Abstract Editor's evaluation Introduction Results Discussion Methods Appendix 1 Data availability References Decision letter Author response Article and author information Metrics Abstract In addition to long-timescale rewiring, synapses in the brain are subject to significant modulation that occurs at faster timescales that endow the brain with additional means of processing information. Despite this, models of the brain like recurrent neural networks (RNNs) often have their weights frozen after training, relying on an internal state stored in neuronal activity to hold task-relevant information. In this work, we study the computational potential and resulting dynamics of a network that relies solely on synapse modulation during inference to process task-relevant information, the multi-plasticity network (MPN). Since the MPN has no recurrent connections, this allows us to study the computational capabilities and dynamical behavior contributed by synapses modulations alone. The generality of the MPN allows for our results to apply to synaptic modulation mechanisms ranging from short-term synaptic plasticity (STSP) to slower modulations such as spike-time dependent plasticity (STDP). We thoroughly examine the neural population dynamics of the MPN trained on integration-based tasks and compare it to known RNN dynamics, finding the two to have fundamentally different attractor structure. We find said differences in dynamics allow the MPN to outperform its RNN counterparts on several neuroscience-relevant tests. Training the MPN across a battery of neuroscience tasks, we find its computational capabilities in such settings is comparable to networks that compute with recurrent connections. Altogether, we believe this work demonstrates the computational possibilities of computing with synaptic modulations and highlights important motifs of these computations so that they can be identified in brain-like systems. Editor's evaluation The study shows that fast and transient modifications of the synaptic efficacies, alone, can support the storage and processing of information over time. Convincing evidence is provided by showing that feed-forward networks, when equipped with such short-term synaptic modulations, perform a wide variety of tasks at a performance level comparable with that of recurrent networks. The results of the study are valuable to both neuroscientists and researchers in machine learning. https://doi.org/10.7554/eLife.83035.sa0 Decision letter Reviews on Sciety eLife's review process Introduction The brain’s synapses constantly change in response to information under several distinct biological mechanisms (Love, 2003; Hebb, 2005; Bailey and Kandel, 1993; Markram et al., 1997; Bi and Poo, 1998; Stevens and Wang, 1995; Markram and Tsodyks, 1996). These changes can serve significantly different purposes and occur at drastically different timescales. Such mechanisms include synaptic rewiring, which modifies the topology of connections between neurons in our brain and can be as fast as minutes to hours. Rewiring is assumed to be the basis of long-term memory that can last a lifetime (Bailey and Kandel, 1993). At faster timescales, individual synapses can have their strength modified (Markram et al., 1997; Bi and Poo, 1998; Stevens and Wang, 1995; Markram and Tsodyks, 1996). These changes can occur over a spectrum of timescales and can be intrinsically transient (Stevens and Wang, 1995; Markram and Tsodyks, 1996). Though such mechanisms may not immediately lead to structural changes, they are thought to be vital to the brain’s function. For example, short-term synaptic plasticity (STSP) can affect synaptic strength on timescales less than a second, with such effects mainly presynaptic-dependent (Stevens and Wang, 1995; Tsodyks and Markram, 1997). At slower timescales, long-term potentiation (LTP) can have effects over minutes to hours or longer, with the early phase being dependent on local signals and the late phase including a more complex dependence on protein synthesis (Baltaci et al., 2019). Also on the slower end, spike-time-dependent plasticity (STDP) adjusts the strengths of connections based on the relative timing of pre- and postsynaptic spikes (Markram et al., 1997; Bi and Poo, 1998; McFarlan et al., 2023). In this work, we investigate a new type of artificial neural network (ANN) that uses biologically motivated synaptic modulations to process short-term sequential information. The multi-plasticity network (MPN) learns using two complementary plasticity mechanisms: (1) long-term synaptic rewiring via standard supervised ANN training and (2) simple synaptic modulations that operate at faster timescales. Unlike many other neural network models with synaptic dynamics (Tsodyks et al., 1998; Mongillo et al., 2008; Lundqvist et al., 2011; Barak and Tsodyks, 2014; Orhan and Ma, 2019; Ballintyn et al., 2019; Masse et al., 2019), the MPN has no recurrent synaptic connections, and thus can only rely on modulations of synaptic strengths to pass short-term information across time. Although both recurrent connections and synaptic modulation are present in the brain, it can be difficult to isolate how each of these affects temporal computation. The MPN thus allows for an in-depth study of the computational power of synaptic modulation alone and how the dynamics behind said computations may differ from networks that rely on recurrence. Having established how modulations alone compute, we believe it will be easier to disentangle synaptic computations from brain-like networks that may compute using a combination of recurrent connections, synaptic dynamics, neuronal dynamics, etc. Biologically, the modulations in the MPN represent a general synapse-specific change of strength on shorter timescales than the structural changes, the latter of which are represented by weight adjustment via backpropagation. We separately consider two forms of modulation mechanisms, one of which is dependent on both the pre- and postsynaptic firing rates and a second that only depends on presynaptic rates. The first of these rules is primarily envisioned as coming from associative forms of plasticity that depend on both pre- and postsynaptic neuron activity (Markram et al., 1997; Bi and Poo, 1998; McFarlan et al., 2023). Meanwhile, the second type of modulation models presynaptic-dependent STSP (Mongillo et al., 2008; Zucker and Regehr, 2002). While both these mechanisms can arise from distinct biological mechanisms and can span timescales of many orders of magnitude, the MPN uses simplified dynamics to keep the effects of synaptic modulations and our subsequent results as general as possible. It is important to note that in the MPN, as in the brain, the mechanisms that represent synaptic modulations and rewiring are not independent of one another – changes in one affect the operation of the other and vice versa. To understand the role of synaptic modulations in computing and how they can change neuronal dynamics, throughout this work we contrast the MPN with recurrent neural networks (RNNs), whose synapses/weights remain fixed after a training period. RNNs store temporal, task-relevant information in transient internal neural activity using recurrent connections and have found widespread success in modeling parts of our brain (Cannon et al., 1983; Ben-Yishai et al., 1995; Seung, 1996; Zhang, 1996; Ermentrout, 1998; Stringer et al., 2002; Xie et al., 2002; Fuhs and Touretzky, 2006; Burak and Fiete, 2009). Although RNNs model the brain’s significant recurrent connections, the weights in these networks neglect the role transient synaptic dynamics can have in adjusting synaptic strengths and processing information. Considerable progress has been made in analyzing brain-like RNNs as population-level dynamical systems, a framework known as neural population dynamics (Vyas et al., 2020). Such studies have revealed a striking universality of the underlying computational scaffold across different types of RNNs and tasks (Maheswaranathan et al., 2019b). To elucidate how computation through synaptic modulations affect neural population behavior, we thoroughly characterize the MPN’s low-dimensional behavior in the neural population dynamics framework (Vyas et al., 2020). Using a novel approach of analyzing the synapse population behavior, we find the MPN computes using completely different dynamics than its RNN counterparts. We then explore the potential benefits behind its distinct dynamics on several neuroscience-relevant tasks. Contributions The primary contributions and findings of this work are as follows: We elucidate the neural population dynamics of the MPN trained on integration-based tasks and show it operates with qualitatively different dynamics and attractor structure than RNNs. We support this with analytical approximations of said dynamics. We show how the MPN’s synaptic modulations allow it to store and update information in its state space using a task-independent, single point-like attractor, with dynamics slower than task-relevant timescales. Despite its simple attractor structure, for integration-based tasks, we show the MPN performs at level comparable or exceeding RNNs on several neuroscience-relevant measures. The MPN is shown to have dynamics that make it a more effective reservoir, less susceptible to catastrophic forgetting, and more flexible to taking in new information than RNN counterparts. We show the MPN is capable of learning more complex tasks, including contextual integration, continuous integration, and 19 neuroscience tasks in the NeuroGym package (Molano-Mazon et al., 2022). For a subset of tasks, we elucidate the changes in dynamics that allow the network to solve them. Related work Networks with synaptic dynamics have been investigated previously (Tsodyks et al., 1998; Mongillo et al., 2008; Sugase-Miyamoto et al., 2008; Lundqvist et al., 2011; Barak and Tsodyks, 2014; Orhan and Ma, 2019; Ballintyn et al., 2019; Masse et al., 2019; Hu et al., 2021; Tyulmankov et al., 2022; Tyulmankov et al., 2022; Rodriguez et al., 2022). As we mention above, many of these works investigate networks with both synaptic dynamics and recurrence (Tsodyks et al., 1998; Mongillo et al., 2008; Lundqvist et al., 2011; Barak and Tsodyks, 2014; Orhan and Ma, 2019; Ballintyn et al., 2019; Masse et al., 2019), whereas here we are interested in investigating the computational capabilities and dynamical behavior of computing with synapse modulations alone. Unlike previous works that examine computation solely through synaptic changes, the MPN’s modulations occur at all times and do not require a special signal to activate their change (Sugase-Miyamoto et al., 2008). The networks examined in this work are most similar to the recently introduced ‘HebbFF’ (Tyulmankov et al., 2022) and ‘STPN’ (Rodriguez et al., 2022) that also examine computation through continuously updated synaptic modulations. Our work differs from these studies in that we focus on elucidating the neural population dynamics of such networks, contrasting them to known RNN dynamics, and show why this difference in dynamics may be beneficial in certain neuroscience-relevant settings. Additionally, the MPN uses a multiplicative modulation mechanism rather than the additive modulation of these two works, which in some settings we investigate yields significant performance differences. The exact form of the synaptic modulation updates were originally inspired by ‘fast weights’ used in machine learning for flexible learning (Ba et al., 2016). However, in the MPN, both plasticity rules apply to the same weights rather than different ones, making it more biologically realistic. This work largely focuses on understanding computation through a neural population dynamics-like analysis (Vyas et al., 2020). In particular, we focus on the dynamics of networks trained on integration-based tasks, that have previously been studied in RNNs (Maheswaranathan et al., 2019b; Maheswaranathan et al., 2019a; Maheswaranathan and Sussillo, 2020; Aitken et al., 2020). These studies have demonstrated a degree of universality of the underlying computational structure across different types of tasks and RNNs (Maheswaranathan et al., 2019b). Due to the MPN’s dynamic weights, its operation is fundamentally different than said recurrent networks. Setup Throughout this work, we primarily investigate the dynamics of the MPN on tasks that require an integration of information over time. To correctly respond to said task, the network is required to both store and update its internal state as well as compare several distinct items in its memory. All tasks in this work consist of a discrete sequence of vector inputs, xt for t=1,2,…,T. For the tasks we consider presently, at time T the network is queried by a ‘go signal’ for an output, for which the correct response can depend on information from the entire input sequence. Throughout this paper, we denote vectors using lowercase bold letters, matrices by uppercase bold letters, and scalars using standard (not-bold) letters. The input, hidden, and output layers of the networks we study have d, n, and N neurons, respectively. Multi-plasticity network The multi-plasticity network (MPN) is an artificial neural network consisting of input, hidden, and output layers of neurons. It is identical to a fully-connected, two-layer, feedforward network (Figure 1, middle), with one major exception: the weights connecting the input and hidden layer are modified by the time-dependent synapse modulation (SM) matrix, M (Figure 1, left). The expression for the hidden layer activity at time step t is (1) ht=tanh⁡((Mt−1⊙Winp)xt+Winpxt) where Winp is an n-by-d weight matrix representing the network’s synaptic strengths that is fixed after training, ‘⊙’ denotes element-wise multiplication of the two matrices (the Hadamard product), and the tanh⁡(⋅) is applied element-wise. For each synaptic weight in Winp, a corresponding element of Mt−1 multiplicatively modulates its strength. Note if Mt−1=0 the first term vanishes, so the Winp are unmodified and the network simply functions as a fully connected feedforward network. Figure 1 Download asset Open asset Two neural network computational mechanisms: synaptic modulations and recurrence. Throughout this figure, neurons are represented as white circles, the black lines between neurons represent regular feedforward weights that are modified during training through gradient descent/backpropagation. From bottom to top are the input, hidden, and output layers, respectively. (Middle) A two-layer, fully connected, feedforward neural network. (Left) Schematic of the MPN. Here, the pink and black lines (between the input and hidden layer) represent weights that are modified by both backpropagation (during training) and the synapse modulation matrix (during an input sequence), see Equation 1. (Right) Schematic of the Vanilla RNN. In addition to regular feedforward weights between layers, the RNN has (fully connected) weights between its hidden layer from one time step to the next, see Equation 3. What allows the MPN to store and manipulate information as the input sequence is passed to the network is how the SM matrix, Mt, changes over time. Throughout this work, we consider two distinct modulation update rules. The primary rule we investigate is dependent upon both the pre- and postsynaptic firing rates. An alternative update rule only depends upon the presynaptic firing rate. Respectively, the SM matrix updated for these two cases takes the form (Hebb, 2005; Ba et al., 2016; Tyulmankov et al., 2022), (2a) pre.&post.:Mt=λMt−1+ηhtxtT (2b) pre. only:Mt=λMt−1+η1xtT/n, where λ and η are parameters learned during training and 1 is the n-dimensional vector of all 1s. We allow for −∞<η<∞, so the size and sign of the modulations can be optimized during training. Additionally, 0<λ<1, so the SM matrix exponentially decays at each time step, asymptotically returning to its M=0 baseline. For both rules, we define M0=0 at the start of each input sequence. Since the SM matrix is updated and passed forward at each time step, we will often refer to Mt as the state of said networks. To distinguish networks with these two modulation rules, we will refer to networks with the presynaptic only rule as MPNpre, while we reserve MPN for networks with the pre- and postsynatpic update that we primarily investigate. For brevity, and since almost all results for the MPN generalize to the simplified update rule of the MPNpre, the main text will foremost focus on results for the MPN. Results for the MPNpre are discussed only briefly or given in the supplement. As mentioned in the introduction, from a biological perspective the MPN’s modulations represent a general associative plasticity such as STDP, whereas the presynaptic-dependent modulations of the MPNpre can represent STSP. The decay induced by λ represents the return to baseline of the aforementioned processes, which all occur at a relatively slow speed to their onset (Bertram et al., 1996; Zucker and Regehr, 2002). To ensure the eventual decay of such modulations, unless otherwise stated, throughout this work we further limit λ<λmax with λmax=0.95. Additionally, we observe no major performance or dynamics difference for positive or negative η, so we do not distinguish the two throughout this work (Methods). We emphasize that the modulation mechanisms of the MPN and MPNpre could represent biological processes that occur at significantly different timescales, so although we train them on identical tasks the tasks themselves are assumed to occur at timescales that match the modulation mechanism of the corresponding network. Note that the modulation mechanisms are not independent of weight adjustment from backpropagation. Since the SM matrix is active during training, the network’s weights that are being adjusted by backpropgation (see below) are experiencing modulations, and said modulations factor into how the weights are adjusted. Lastly, the output of the MPN and MPNpre at time T is determined by a fully-connected readout matrix, yT=WROhT, where WRO is an N-by-n weight matrix adjusted during training. Throughout this work, we will view said readout matrix as N distinct n-dimensional readout vectors, that is one for each output neuron. Recurrent neural networks As discussed in the introduction, throughout this work we will compare the learned dynamics and performance of the MPN to artificial RNNs. The hidden layer activity for the simplest recurrent neural network, the Vanilla RNN, is (3) ht=tanh⁡(Wrecht−1+Winpxt+b), with Wrec the recurrent weights, an n-by-n matrix that updates the hidden neurons from one time step to the next (Figure 1, right). We also consider a more sophisticated RNN structure, the gated recurrent unit (GRU), that has additional gates to more precisely control the recurrent update of its hidden neurons (see Methods 5.2). In both these RNNs, information is stored and updated via the hidden neuron activity, so we will often refer to ht as the RNNs’ hidden state or just its state. The output of the RNNs is determined through a trained readout matrix in the same manner as the MPN above, i.e. yT=WROhT. Training The weights of the MPN, MPNpre, and RNNs will be trained using gradient descent/backpropagation through time, specifically ADAM (Kingma and Ba, 2014). All network weights are subject to L1 regularization to encourage sparse solutions (Methods 5.2). Cross-entropy loss is used as a measure of performance during training. Gaussian noise is added to all inputs of the networks we investigate. Results Network dynamics on a simple integration task Simple integration task We begin our investigation of the MPN’s dynamics by training it on a simple N-class (Through most of this work, the number of neurons in the output layer of our networks will always be equal to the number of classes in the task, so we use N to denote both unless otherwise stated). integration task, inspired by previous works on RNN integration-dynamics (Maheswaranathan et al., 2019a; Aitken et al., 2020). In this task, the network will need to determine for which of the N classes the input sequence contains the most evidence (Figure 2a). Each stimulus input, xt, can correspond to a discrete unit of evidence for one of the N classes. We also allow inputs that are evidence for none of the classes. The final input, xT, will always be a special ‘go signal’ input that tells the network an output is expected. The network’s output should be an integration of evidence over the entire input sequence, with an output activity that is largest from the neuron that corresponds to the class with the maximal accumulated evidence. (We omit sequences with two or more classes tied for the most evidence. See Methods 5.1 for additional details). Prior to adding noise, each possible input, including the go signal, is mapped to a random binary vector (Figure 2b). We will also investigate the effect of inserting a delay period between the stimulus period and the go signal, during which no input is passed to the network, other than noise (Figure 2c). Figure 2 Download asset Open asset Schematic of simple integration task. (a) Example sequence of the two-class integration task where each represents an and throughout this work, distinct classes are represented by different In this and The represent evidence for their while the represents an input that is evidence for At the of the sequence is the ‘go signal’ that the network an output is expected. The correct response for the sequence is the class with the most in the the Each possible input is mapped to a random binary The integration task can be modified by the of a between the stimulus period and the go the delay the network no input than We find the MPN is capable of learning the integration task to across a wide of class sequence and delay It is the of this to the dynamics behind the trained MPN that allow it to solve such a task and compare them to more RNN dynamics. Here, in the main we will explore the dynamics of a two-class integration task, to classes are and are discussed in the Methods We will start by the simplest of integration a delay the effects of delay we into the dynamics of the MPN, we a of the known RNN dynamics on integration-based tasks. of RNN attractor dynamics accumulated evidence both on and artificial neural networks, have that networks with recurrent connections attractor dynamics to solve integration-based tasks (Maheswaranathan et al., 2019a; Maheswaranathan and Sussillo, 2020; Aitken et al., 2020). Here, we specifically review the behavior of artificial RNNs on the aforementioned N-class integration tasks that many with of neural networks. Note also the of the dynamics can depend on between the classes et al., in this work we only investigate the where the classes are RNNs are capable of learning to solve the simple integration task at and their dynamics are qualitatively the same across several (Maheswaranathan et al., 2019b; Aitken et al., 2020). the network’s behavior by at individual hidden neuron activity can be difficult (Figure and so it is to to a population-level analysis of the dynamics. the number of hidden neurons is than number of integration classes the population activity of the trained RNN primarily in a low-dimensional of et al., 2020). This is to recurrent dynamics that a attractor of and the hidden activity often operates to said Methods for a more in-depth review of these results including how is In the two-class the RNN will operate to a The of hidden activity allows for an of the dynamics using a (Figure From the we see the network’s hidden activity from the attractor its As evidence for one class over the other the hidden activity accumulated evidence by the attractor (Figure The two readout vectors are with the two of the so the further the final hidden activity, is one of the attractor, the that corresponding output and thus the RNN correctly the class with the most evidence. For we note that the hidden activity of the trained RNN is not dependent upon the input of the present time step (Figure it is the change in the hidden activity from one time step to the next, that are (Figure For the Vanilla RNN (GRU), we find of the hidden activity to be by the accumulated evidence and only to be by the present input to the network Methods Figure with see all Download asset Open asset of multi-plasticity network and RNN dynamics. Vanilla RNN hidden neuron dynamics, see Figure 1 for (a) layer neural activity, for neurons of the RNN as a of sequence time of sequence The represents the stimulus period during which information should be across time and the representing the response to the go neuron activity, over input into their top two by relative accumulated evidence between classes at time t (Methods Also shown are of by the class readout vector and the state as with ht by input at the present time step, xt see The shows the of as a of the present input, xt, with the lines showing the for each of the MPN hidden neuron dynamics, see Figure 1 for as as as for The shows the of each with the readout vectors (Methods MPN synaptic modulation dynamics. as of hidden neuron activity, the of the SM Mt, over input Mt are for as with a different The is the same as that shown in for MPN hidden activity inputs, not so accumulated evidence We to analyzing the hidden activity of the trained in the same manner that for the RNNs. The MPN trained on a two-class integration task to have significantly more activity in the individual of ht (Figure We find the hidden neuron activity to be with it to using a (Methods Unlike the RNN, we observe the hidden neuron activity to be into several distinct (Figure input sequences ht to between said the ht by the sequence input at the present time step, we see the different inputs are the hidden activity into distinct that we (Figure the hidden neuron activity is largely dependent upon the most input to the network, rather than the accumulated evidence as we for the RNN. However, each we also see a in ht from accumulated evidence (Figure For the MPN we find only of the hidden activity to be by accumulated evidence and to be by the present input to the network Methods MPNpre dynamics are largely the same of we see for further the hidden neuron activity primarily dependent upon the input to the network, one may how the MPN information dependent upon the entire sequence to solve the task. the other possible inputs to the network, the go signal has its distinct which the hidden by accumulated evidence. all we find the readout vectors are with the evidence the go (Figure The are

  • Peer Review Report
  • 10.7554/elife.83035.sa1
Decision letter: Neural population dynamics of computing with synaptic modulations
  • Jan 8, 2023
  • Omri Barak

Synaptic modulations alone imbue networks with computational capabilities comparable to recurrent connections on several neuroscience-relevant tasks, which manifest in fundamentally different neuronal dynamics.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 52
  • 10.1038/s41598-024-77687-x
Predicting rainfall using machine learning, deep learning, and time series models across an altitudinal gradient in the North-Western Himalayas
  • Nov 13, 2024
  • Scientific Reports
  • Owais Ali Wani + 8 more

Predicting rainfall is a challenging and critical task due to its significant impact on society. Timely and accurate predictions are essential for minimizing human and financial losses. The dependence of approximately 60% of agricultural land in India on monsoon rainfall implies the crucial nature of accurate rainfall prediction. Precise rainfall forecasts can facilitate early preparedness for disasters associated with heavy rains, enabling the public and government to take necessary precautions. In the North-Western Himalayas, where meteorological data are limited, the need for improved accuracy in traditional modeling methods for rainfall forecasting is pressing. To address this, our study proposes the application of advanced machine learning (ML) algorithms, including random forest (RF), support vector regression (SVR), artificial neural network (ANN), and k-nearest neighbour (KNN) along with various deep learning (DL) algorithms such as long short-term memory (LSTM), bi-directional LSTM, deep LSTM, gated recurrent unit (GRU), and simple recurrent neural network (RNN). These advanced techniques hold the potential to significantly improve the accuracy of rainfall prediction, offering hope for more reliable forecasts. Additionally, time series techniques, including autoregressive integrated moving average (ARIMA) and trigonometric, Box-Cox transform, arma errors, trend, and seasonal components (TBATS), are proposed for predicting rainfall across the altitudinal gradients of India’s North-Western Himalayas. This approach can potentially revolutionise how we approach rainfall forecasting, ushering in a new era of accuracy and reliability. The effectiveness and accuracy of the proposed algorithms were assessed using meteorological data obtained from six weather stations at different elevations spanning from 1980 to 2021. The results indicate that DL methods exhibit the highest accuracy in predicting rainfall, as measured by the root mean squared error (RMSE) and mean absolute error (MAE), followed by ML algorithms and time series techniques. Among the DL algorithms, the accuracy order was bi-directional LSTM, LSTM, RNN, deep LSTM, and GRU. For the ML algorithms, the accuracy order was ANN, KNN, SVR, and RF. These findings suggest that altitude significantly affects the accuracy of the models, highlighting the need for additional weather stations in this mountainous region to enhance the precision of rainfall prediction.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.heliyon.2024.e36134
TrajPredRNN+: A new approach for precipitation nowcasting with weather radar echo images based on deep learning
  • Aug 10, 2024
  • Heliyon
  • Chongxing Ji + 1 more

trajPredRNN+: A new approach for precipitation nowcasting with weather radar echo images based on deep learning

  • Conference Article
  • Cite Count Icon 28
  • 10.1109/icphm.2019.8819440
Deep Recurrent Convolutional Neural Network for Remaining Useful Life Prediction
  • Jun 1, 2019
  • Meng Ma + 1 more

Remaining Useful Life (RUL) prediction of rotating machinery plays a critical role in Prognostics and Health Management (PHM). Data-driven methods for RUL estimation have been widely developed because they don’t depend on much prior knowledge of the system. Recurrent neural network (RNN) is capable of modeling sequential data, which has been investigated for RUL prediction with statistical features of vibration signals in time domain and frequency domain. The drawback of utilizing statistical features is the ignorance of time-frequency information, which is critical in RUL prediction because the vibration signals are non-stationary when the fault occurs. To solve this problem, a novel deep architecture, named deep recurrent convolutional neural network (DRCNN) is proposed. By incorporating convolutional operation in the process of state transition of RNN, the spatial information in time-frequency domain can be automatically learned from the vibration signals, which contributes to the improvement of prediction performance. With convolutional operation in RNN, both spatial information in time-frequency domain and previous information are employed for RUL prediction. Furthermore, by stacking recurrent convolutional neural network layer by layer, the deep architecture can learn high-level features in the time-frequency domain. Finally, experimental analysis of RUL prediction using vibration signals of run-to-failure tests are carried out. Compared with the results of conventional deep RNN method, the proposed method shows its effectiveness and superiority.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant