Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

CryptoMamba-SSM: Linear Complexity State Space Models for Cryptocurrency Volatility Prediction

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

CryptoMamba-SSM introduces a linear-complexity state space model for cryptocurrency volatility prediction, effectively capturing regime shifts and microstructural signals. It outperforms LSTM, GRU, and Transformer models with up to 23.7% lower MAE and 31.2% higher directional accuracy, enabling efficient real-time analysis across market regimes.

Abstract
Translate article icon Translate Article Star icon

Cryptocurrency markets exhibit complex microstructural dynamics characterized by high-frequency volatility bursts, rapid regime switching, and long-range temporal dependencies, which expose several limitations of existing volatility forecasting approaches. In particular, attention-based models suffer from prohibitive quadratic computational cost on long high-frequency sequences, while many recurrent architectures struggle to adapt to regime transitions, asymmetric volatility responses, and risk-aware uncertainty estimation. To address these gaps, this paper proposes <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CryptoMamba-SSM</b>, a novel volatility prediction framework built upon Mamba-based state space models with linear computational complexity. CryptoMamba-SSM integrates selective memory mechanisms with structured state space representations to effectively capture critical market microstructure signals arising from liquidity shocks and sentiment transitions, while dynamically adjusting memory retention across different volatility regimes. This design enables efficient modeling of long-sequence dependencies inherent in cryptocurrency price movements without incurring the computational bottlenecks of traditional attention-based architectures. Through comprehensive experiments on Bitcoin historical data spanning multiple market regimes, we demonstrate that CryptoMamba-SSM consistently outperforms conventional LSTM, GRU, and Transformer baselines, achieving up to a 23.7% reduction in Mean Absolute Error and a 31.2% improvement in directional accuracy. The selective memory mechanism effectively captures regime-switching behaviors and microstructural anomalies, leading to more reliable short-term volatility risk quantification. Moreover, the linear-time complexity of CryptoMamba-SSM enables real-time processing of high-frequency trading data while maintaining strong generalization across diverse market conditions.

Similar Papers
  • Research Article
  • Cite Count Icon 29
  • 10.1007/s10489-012-0417-1
Comparison of individual and combined ANN models for prediction of air and dew point temperature
  • Feb 3, 2013
  • Applied Intelligence
  • Karthik Nadig + 3 more

Predicted air and dew point temperatures can be valuable in decision making in many areas including protecting crops from damage, avoiding heat stress on animals and humans, and in planning related to energy management. Current web-based artificial neural network (ANN) models on the Automated Environment Monitoring Network (AEMN) in Georgia predict hourly air and dew point temperature for twelve prediction horizons, using 24 models. The observed air temperature may approach the observed dew point temperature, but never goes below it. Current web based ANN models have prediction errors which, when the air and dew point temperatures are close, may cause air temperature to be predicted below the dew point temperature. Herein this error is referred to as a prediction anomaly. The goal of this research was to improve the prediction accuracy of existing air and dew point temperature ANN models by combining the two weather variables into a single ANN model for each prediction horizon. The objectives of this study were to reduce the mean absolute error (MAE) of prediction and to reduce the number of prediction anomalies. The combined models produced a reduction in the air temperature MAE for ten of twelve prediction horizons with an average reduction in MAE of 1.93 %. The combined models produced a reduction in the dew point temperature MAE for only six of twelve prediction horizons with essentially no average decrease in MAE. However, the combined models showed a marked reduction in prediction anomalies for all twelve prediction horizons with an average reduction of 34.1 %. The reduction in prediction anomalies ranged from 4.6 % at the one-hour horizon to 60.5 % at the eleven-hour horizon.

  • Research Article
  • Cite Count Icon 38
  • 10.1016/j.eswa.2024.123697
Forecasting price in a new hybrid neural network model with machine learning
  • Mar 18, 2024
  • Expert Systems with Applications
  • Rui Zhu + 2 more

Forecasting price in a new hybrid neural network model with machine learning

  • Research Article
  • 10.1186/s12859-026-06383-6
Transplatformer: translating toxicogenomic profiles between generations of platforms.
  • Jan 30, 2026
  • BMC bioinformatics
  • Guojing Cong + 9 more

Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation (≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra CONCLUSIONS: TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.5194/gi-11-389-2022
Upgrade of LSA-SAF Meteosat Second Generation daily surface albedo (MDAL) retrieval algorithm incorporating aerosol correction and other improvements
  • Nov 24, 2022
  • Geoscientific Instrumentation, Methods and Data Systems
  • Daniel Juncu + 4 more

Abstract. MDAL is the operational Meteosat Second Generation (MSG)-derived daily surface albedo product that has been generated and disseminated in near real time by EUMETSAT Satellite Application Facility for Land Surface Analysis (LSA-SAF) since 2005. We propose and evaluate an update to the MDAL retrieval algorithm which introduces the accounting for aerosol effects as well as other scientific developments: pre-processing recalibration of radiances acquired by the SEVIRI instrument aboard MSG and improved coefficients for atmospheric correction as well as for albedo conversion from narrow- to broadband. We compare the performance of MDAL broadband albedos pre- and post-upgrade with respect to three types of reference data: the EPS Ten-Day Albedo product ETAL is used as the primary reference, while albedo derived from in situ flux measurements acquired by ground stations and MODIS MCD43D albedo data are used to complete the validation. For the comparison to ETAL – conducted over the whole coverage area of SEVIRI – we see a reduction in average white-sky albedo mean bias error (MBE) from −0.02 to negligible levels (&lt;0.001) and a reduction in average mean absolute error (MAE) from 0.034 to 0.026 (−24 %). Improvements can be seen for black-sky albedo as well, albeit less pronounced (14 % reduction in MAE). Further analysis distinguishing individual seasons, regions and land covers show that performance changes have spatial and temporal dependence: for white-sky albedo we see improvements over almost all regions and seasons relative to ETAL, except for Eurasia in winter; resolved by land cover we see a similar effect with improvements for all types for all seasons except winter, where some types exhibit slightly worse results (crop-, grass- and shrublands). For black-sky albedo we similarly see improvements for all seasons when averaged over the full data set, although sub-regions exhibit clear seasonal dependence: the performance of the upgraded MDAL version is generally diminished in local winter but better in local summer. The comparison with in situ observations is less conclusive due to the well-known problem of the spatial representativeness of near-ground observations with respect to satellite pixel footprint sizes. Comparison with MODIS at the same locations shows mixed results in terms of change in performance following the proposed upgrade but proves the good quality of the MDAL products in general. Based on the evidence presented in this study, we consider the updated algorithm version to be able to deliver a valuable improvement of the operational MDAL product. This improvement is two-fold: primarily, there is the refinement of the albedo values themselves; secondarily, the increased alignment with the ETAL product is beneficial for those who wish to exploit synergies between EUMETSAT's geostationary and polar satellites to generate data sets based on the LSA-SAF albedo products from the two different missions.

  • Book Chapter
  • Cite Count Icon 12
  • 10.1007/978-1-4612-0761-0_5
Linear Gaussian State Space Modeling
  • Jan 1, 1996
  • Genshiro Kitagawa + 1 more

Linear Gaussian state space modeling is treated in this chapter. The prediction, filtering and smoothing formulas in the standard Kalman filter are shown. Model identification or, computation of the likelihood of the model is also treated. Some of the well known state space models that are used in this book as well as state space modeling of missing observations and a state space model for unequally spaced time series are shown. The final section is a discussion of the information square root filter/smoother, that we use in linear Gaussian state space seasonal decomposition modeling in Chapter 9. Not necessarily linear - not necessarily Gaussian state space modeling is treated in Chapter 6. A variety of illustrative examples of linear state space modeling is shown in Chapter 7.

  • Research Article
  • Cite Count Icon 8
  • 10.1016/j.jbi.2021.103975
Improving non-invasive hemoglobin measurement accuracy using nonparametric models
  • Dec 11, 2021
  • Journal of Biomedical Informatics
  • Jianing Man + 4 more

Improving non-invasive hemoglobin measurement accuracy using nonparametric models

  • Research Article
  • Cite Count Icon 1
  • 10.1109/access.2025.3603070
Optimized Transformer and GRU Models for Forecasting Lost Circulation Volume in Drilling Operations
  • Jan 1, 2025
  • IEEE Access
  • Bayan Zumarah + 3 more

Lost circulation remains a persistent challenge in drilling operations, often resulting in increased non-productive time (NPT), elevated operational costs, and heightened well control risks. While prior research has primarily focused on the detection or classification of mud losses, this study addresses the short-term forecasting of the lost circulation volume (LCV) using advanced deep learning techniques. We investigated two modeling approaches. First, a series of deep recurrent neural network (RNN) models based on Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and hybrid configurations were developed to predict the LCV (in barrels) for the next short time interval. Second, a Transformer-based approach was introduced, with nine models exploring various architectural configurations. In total, twenty-two models were developed, including thirteen RNN-based models and 9 Transformer-based models. The dataset was divided into training, testing, and validation sets. A final validation set was held out from the initial split to simulate unseen conditions. The models were evaluated using standard regression metrics, with a particular focus on the coefficient of determination (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>) and the Mean Absolute Error (MAE). Our Transformer-based model outperforms traditional ML benchmarks, achieving up to 3.21% higher <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> and reducing MAE by approximately 36.3%. Specifically, for the 15-minute aggregated testing data, our model achieves a 2.69% improvement in <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> and a 33.5% reduction in MAE compared to the Random Forest model. For the 30-minute aggregated testing data, our model achieves a 2.99% improvement in <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> and a 35.1% reduction in MAE compared to the Gradient Boosting model. For the 60-minute aggregated testing data, our model achieves a 3.21% improvement in <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> and a 36.3% reduction in MAE compared to the MLP model. The best-performing model achieved <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</i><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> scores of 0.9882, 0.9910, and 0.9932 on the 15-minute, 30-minute, and 60-minute aggregated testing data, respectively. Corresponding MAE scores were 0.0243, 0.0459, and 0.0837 barrels. These findings highlight the effectiveness of deep learning, especially Transformer architectures, for accurate and timely forecasting of LCV. This provides a practical tool for early loss detection, proactive fluid management, and improved operational efficiency in drilling workflows.

  • Research Article
  • Cite Count Icon 3
  • 10.1080/19942060.2025.2538180
Investigation of the impact of token embeddings in Transformer-based models on short-term tropical cyclone track and intensity predictions
  • Aug 1, 2025
  • Engineering Applications of Computational Fluid Mechanics
  • Yuan-Jiang Zeng + 5 more

Tropical cyclones (TCs) are destructive meteorological phenomena, necessitating accurate predictions of TC track and intensity to reduce risks to human life. This study evaluates three Transformer-based models – vanilla Transformer (Transformer), inverted Transformer (iTransformer), and temporal-variate Transformer (TVFormer) – which are trained, validated, and tested on best track data from 1980 to 2021 from the China Meteorological Administration for TC prediction, integrating temporal, variate, and hybrid token embeddings to analyze temporal and variate correlations. Comparative analysis with four recurrent neural network (RNN) models demonstrates the superiority of the refined Transformer models over RNNs: iTransformer reduces mean absolute error (MAE) and root mean square error (RMSE) by 29.55% and 25.80% (latitude), 50.31% and 46.18% (longitude), 8.71% and 9.98% (pressure), and 8.68% and 9.45% (wind speed), while TVFormer achieves MAE and RMSE reductions of 13.98% and 13.84% (latitude), 39.11% and 38.02% (longitude), 13.69% and 14.02% (pressure), and 12.84% and 12.94% (wind speed) on average. Among Transformer variants, iTransformer excels in track prediction, outperforming Transformer with 21.74% lower MAE and 18.26% lower RMSE for latitude, and 32.73% lower MAE and 24.01% lower RMSE for longitude. TVFormer dominates intensity prediction, reducing pressure errors by 4.42% (MAE) and 3.92% (RMSE) and wind speed errors by 19.21% (MAE) and 14.79% (RMSE) compared to Transformer, while outperforming iTransformer with 4.59% lower MAE and 3.68% lower RMSE for pressure and 3.83% lower MAE and 3.18% lower RMSE for wind speed. Notably, TVFormer also enhances track prediction, with 7.10% reduction in MAE and 7.02% reduction in RMSE for latitude, and 22.84% reduction in MAE and 17.09% reduction in RMSE for longitude compared to Transformer. These results highlight the superiority of iTransformer in track prediction and the efficacy of TVFormer in intensity prediction, thanks to their ability to exploit temporal and variate dependencies, offering potential for TC disaster preparedness systems.

  • Research Article
  • Cite Count Icon 9
  • 10.1016/j.energy.2013.08.052
State space model extraction of thermohydraulic systems – Part II: A linear graph approach applied to a Brayton cycle-based power conversion unit
  • Sep 23, 2013
  • Energy
  • Kenneth Richard Uren + 1 more

State space model extraction of thermohydraulic systems – Part II: A linear graph approach applied to a Brayton cycle-based power conversion unit

  • Research Article
  • Cite Count Icon 10
  • 10.1155/2024/6305525
A Hybrid GARCH and Deep Learning Method for Volatility Prediction
  • Jan 1, 2024
  • Journal of Applied Mathematics
  • Hailabe T Araya + 2 more

Volatility prediction plays a vital role in financial data. The time series movements of stock prices are commonly characterized as highly nonlinear and volatile. This study is aimed at enhancing the accuracy of return volatility forecasts for stock prices by investigating the prediction of their price volatility through the integration of diverse models. Thus, the study integrated four powerful methods: seasonal autoregressive (AR) integrated moving average (MA), generalized AR conditional heteroskedasticity (ARCH) family models, convolutional neural network (CNN), and bidirectional long short‐term memory (LSTM) network. The hybrid model was developed using the residuals generated by the seasonal AR integrated MA model as input for the generalized ARCH model. Following this, the estimated volatility obtained was utilized as an input feature for both the hybrid CNNs and bidirectional LSTM models. The model’s forecasting performance was assessed using key evaluation metrics, including mean absolute error (MAE) and root mean squared error (RMSE). Compared to other hybrid models, our new proposed hybrid model demonstrates an average reduction in MAE and RMSE of 60.35% and 60.61%, respectively. The experimental results show that the model proposed in this study has good performance and accuracy in predicting the volatility of stock prices. These findings offer valuable insights for financial data analysis and risk management strategies.

  • Research Article
  • 10.1049/joe.2018.9048
Variational Bayesian inference of linear state space models
  • Nov 29, 2019
  • The Journal of Engineering
  • Chuanchao Pan + 2 more

This article studies a variational Bayesian method to fix the linear regression (LR) model of which regressors are Gaussian distributed with non-zero prior means, and then apply the method to the linear state space (LSS) model. Here, we innovatively transform the LSS model into a special LR model: In each state, the value obtained from the predict step can be seen as the prior mean of the regressors, and the update step can be viewed as the iterative solving in LR model with non-zero prior means. We simulate the proposed algorithm with high-dimensional discrete LSS models where most states are prior zeros; simulation results show that the proposed algorithm and its applications in LSS are both effective and reliable.

  • Research Article
  • 10.21037/qims-2025-480
Automated method for quantitative analysis of iris fluorescein angiography based on machine learning
  • Dec 31, 2025
  • Quantitative Imaging in Medicine and Surgery
  • Yixuan Zhu + 5 more

BackgroundDiabetic retinopathy is a leading cause of vision impairment, often progressing to neovascular glaucoma. Early detection of neovascularisation of the iris (NVI) is crucial for timely intervention. Traditional diagnostic methods, such as slit-lamp examination, have limitations in identifying early-stage NVI. This study presents a deep learning-based automated approach for analysing iris fluorescein angiography (IFA) images to detect and quantify peripupillary leakage, a key indicator of NVI.MethodsA dataset of 2,449 IFA images was used to train a YOLOv8n-based segmentation model for precise pupil localisation. A leakage circularity detection algorithm was developed to quantify peripupillary fluorescein leakage. The algorithm’s performance was evaluated using an independent test set of 131 clinically standardized IFA images. Performance metrics included mean absolute error (MAE), mean absolute percentage error (MAPE), and intersection over union (IoU). Results were compared with manual annotations from two clinical experts.ResultsThe proposed method demonstrated a significant reduction in MAE (20.81 degrees) and MAPE (21.64%) compared to Clinical Staff 1 (MAE: 34.23 degrees, MAPE: 58.38%) and Clinical Staff 2 (MAE: 43.17 degrees, MAPE: 75.71%). The algorithm achieved an IoU of 39.3%, slightly lower than Clinical Staff 1 (44.5%) and Clinical Staff 2 (41.7%), indicating high segmentation accuracy but minor spatial misalignment. The inter-clinician agreement yielded an IoU of 54.8%, highlighting subjectivity in human assessments.ConclusionsThe deep learning-based approach provides superior consistency and accuracy in quantifying peripupillary fluorescein leakage compared to manual expert annotations. While human experts demonstrated slightly higher spatial precision, the algorithm significantly reduces variability and subjectivity in leakage quantification. This automated method has the potential to enhance early detection of NVI, improve clinical workflow efficiency, and assist ophthalmologists in diagnosing DR. Further optimization will focus on refining spatial segmentation accuracy.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/icmlc.2012.6359017
The hourly load forecasting based on linear Gaussian state space model
  • Jul 1, 2012
  • Yanxia-Lu + 1 more

In this paper, the linear gaussian state space model is used to forecast the hourly electricity load. Since the weather variables have significant impacts on electricity demand, thus in our forecasting model, the weather variables are considered as explanatory variables and added to the state space model. The variance parameters of the linear gaussian state space are estimated by the Markov chain Monte Carlo method. Given the estimated parameters, the linear gaussian state space is used to forecast the electricity load on two hours SAM and 14PM respectively. The result shows that this model has higher forecasting precision than the one to four days ahead forecasting, and the state space model estimated by Gibbs sampling algorithm has better performance than the model based on the MH algorithm.

  • Research Article
  • 10.1364/oe.584154
Temporal deep neural network for tactile sensing in artificial finger pulp skin.
  • Feb 9, 2026
  • Optics express
  • Zhiyuan Xu + 6 more

Tactile perception plays a vital role in artificial finger pulp skin, especially in regions responsible for grasping and touching tasks, where precise sensing of deformation position and applied force is critical. Conventional demodulation methods often fail to fully leverage temporal correlations in data during the pressing process, limiting the accuracy of tactile demodulation. To address this, we propose a tactile sensing system based on quasi-distributed Fiber Bragg Gratings (FBGs) integrated into artificial finger pulp skin, along with a two-stage hybrid LSTM-Transformer neural network (TSH-LTNN) to jointly reconstruct pressing position and force. The network trains a temporal demodulation model by constructing possible data variations over three consecutive time steps, where the LSTM captures short-term continuity, the Transformer extracts long-range dependencies, and an adaptive fusion module integrates their complementary features. Experimental results show that the proposed model outperforms existing methods. In the 0-30 mm pressing range and 0-14.71 N force range, the mean absolute error (MAE) for position prediction is 0.2331 mm (R2 = 0.9971), and for force prediction, it is 0.303 N (R2 = 0.9829). Compared to the Random Forest model, the TSH-LTNN achieves a 2.34% improvement in position R2 and a 66.06% reduction in MAE. For force prediction, it demonstrates a 2.48% improvement in R2 and a 31.94% reduction in MAE. These results confirm that the proposed system offers precise, stable, and real-time pressure-state demodulation, with strong potential for high-precision haptic feedback applications.

  • Research Article
  • Cite Count Icon 1
  • 10.1177/03611981231189733
Enhanced Crash Frequency Models Using Surrogate Safety Measures from Connected Vehicle Fleet
  • Jul 31, 2023
  • Transportation Research Record: Journal of the Transportation Research Board
  • Taehun Lee + 1 more

As connected vehicle data became available, efforts to employ surrogate safety measures (SSMs) for crash frequency modeling were undertaken. For road safety evaluation, traffic conflicts are quantitatively measured by various SSMs in two dimensions: spatial/temporal proximity (e.g., time-to-collision, TTC) and evasive action (e.g., deceleration rate, DR). However, a single SSM or a single dimension only represents partial images of the true severity of traffic conflicts. Therefore, this study investigates possible enhancements in crash frequency modeling by concurrently using proximity and evasion SSMs. For rear-end crash frequency estimation, five negative binomial regression models and two tree-based models (a regression tree and random forest) were developed. All models were estimated using crashes, traffic volume, and segment length, along with three SSMs (DR, TTC, and modified TTC) extracted from connected vehicle data in Ann Arbor, Michigan. Results show that the multi-SSM model produced a 19.3% reduction in mean absolute error (MAE) compared to the baseline model with no SSM variable, which was significantly higher than those of single-SSM models (4.5−6.9% reductions in MAE). Between all models, the random forest, which is the ensemble machine learning model, produced the highest error reductions (a 44.3% reduction in MAE). These findings show that the concurrent use of proximity and evasion SSMs can yield further enhancements in crash frequency models compared to the singular use of either type of SSM. The proposed modeling method can be used for proactive safety management and assessment using connected vehicle data collected over a short period.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant