Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Integrating Multi-Scale Acoustic Features with State Space Model for Speech Separation

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Integrating Multi-Scale Acoustic Features with State Space Model for Speech Separation

Similar Papers
  • Research Article
  • Cite Count Icon 2
  • 10.1109/tgrs.2025.3616967
CWIMamba: Cross-Scale Windowed Integration State Space Model for Hyperspectral Anomaly Detection
  • Jan 1, 2025
  • IEEE Transactions on Geoscience and Remote Sensing
  • Xu He + 6 more

Hyperspectral anomaly detection (HAD) intends to detect potential anomalous targets hidden in the background of hyperspectral images (HSIs) and has garnered substantial attention in various remote sensing photography and surveying applications. Recent research advances in the HAD domain have highlighted the significance of deep convolutional networks (DCNs) and vision transformers (ViTs)-based formulas. However, DCNs are long-range dependency-limited networks, whereas ViTs bear the computational burden of quadratic complexity. Owing to their prominent nonlocal representations and linear complexity, Mamba-based approaches have drawn growing attention. Our study pioneers the integration of Mamba into HAD tasks, presenting CWIMamba, which introduces a novel cross-scale windowed integration state space model for considering the spatial distribution characteristics of the anomaly targets. Specifically, we devise a cross-scale windowed state space model (CSWSSM) to scan the spatial-spectral features based on the window-based bottleneck SSM with different scales. For better multiscale feature integration, a multiscale spatial-spectral feature adaptive integration (MS<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup>FAI) method is explored to generate an intensified representation of multiscale feature interaction and fusion based on the elaborate adaptive spatial-spectral weighting scheme. Moreover, we also devised a Haar discrete wavelet transform convolution module (HDWTCM) to fully replenish the local informative representation and enhance the discriminative frequency characteristics between anomalies and background, introducing more inductive local features for accurate background reconstruction and anomaly suppression. Extensive experiments on five multifarious HAD datasets and seven indicators substantiate the state-of-the-art detection performance, demonstrating the effectiveness of CWIMamba.

  • Research Article
  • 10.3390/electronics14193874
FFMamba: Feature Fusion State Space Model Based on Sound Event Localization and Detection
  • Sep 29, 2025
  • Electronics
  • Yibo Li + 3 more

Previous studies on Sound Event Localization and Detection (SELD) have primarily focused on CNN- and Transformer-based designs. While CNNs possess local receptive fields, making it difficult to capture global dependencies over long sequences, Transformers excel at modeling long-range dependencies but have limited sensitivity to local time–frequency features. Recently, the VMamba architecture, built upon the Visual State Space (VSS) model, has shown great promise in handling long sequences, yet it remains limited in modeling local spatial details. To address this issue, we propose a novel state space model with an attention-enhanced feature fusion mechanism, termed FFMamba, which balances both local spatial modeling and long-range dependency capture. At a fine-grained level, we design two key modules: the Multi-Scale Fusion Visual State Space (MSFVSS) module and the Wavelet Transform-Enhanced Downsampling (WTED) module. Specifically, the MSFVSS module integrates a Multi-Scale Fusion (MSF) component into the VSS framework, enhancing its ability to capture both long-range temporal dependencies and detailed local spatial information. Meanwhile, the WTED module employs a dual-branch design to fuse spatial and frequency domain features, improving the richness of feature representations. Comparative experiments were conducted on the DCASE2021 Task 3 and DCASE2022 Task 3 datasets. The results demonstrate that the proposed FFMamba model outperforms recent approaches in capturing long-range temporal dependencies and effectively integrating multi-scale audio features. In addition, ablation studies confirmed the effectiveness of the MSFVSS and WTED modules.

  • Research Article
  • 10.3390/s25195967
VSM-UNet: A Visual State Space Reconstruction Network for Anomaly Detection of Catenary Support Components
  • Sep 25, 2025
  • Sensors (Basel, Switzerland)
  • Shuai Xu + 5 more

Anomaly detection of catenary support components (CSCs) is an important component in railway condition monitoring systems. However, because the abnormal features of CSCs loosening are not obvious, and the current CNN models and visual Transformer models have problems such as limited remote modeling capabilities and secondary computational complexity, it is difficult for existing deep learning anomaly detection methods to effectively exert their performance. The state space model (SSM) represented by Mamba is not only good at long-range modeling, but also maintains linear computational complexity. In this paper, using the state space model (SSM), we proposed a new visual state space reconstruction network (VSM-UNet) for the detection of CSC loosening anomalies. First, based on the structure of UNet, a visual state space block (VSS block) is introduced to capture extensive contextual information and multi-scale features, and an asymmetric encoder–decoder structure is constructed through patch merging operations and patch expanding operations. Secondly, the CBAM attention mechanism is introduced between the encoder–decoder structure to enhance the model’s ability to focus on key abnormal features. Finally, a stable abnormality score calculation module is designed using MLP to evaluate the degree of abnormality of components. The experiment shows that the VSM-UNet model, learning strategy and anomaly score calculation method proposed in this article are effective and reasonable, and have certain advantages. Specifically, the proposed method framework can achieve an AUROC of 0.986 and an FPS of 26.56 in the anomaly detection task of looseness on positioning clamp nuts, U-shaped hoop nuts, and cotton pins. Therefore, the method proposed in this article can be effectively applied to the detection of CSCs abnormalities.

  • Research Article
  • Cite Count Icon 1
  • 10.1371/journal.pone.0330740
VM-Unet enhanced with multi-scale pyramid feature extraction for segmentation of tibiofemoral joint tissues from knee MRI.
  • Aug 28, 2025
  • PloS one
  • Xin Wang + 4 more

In medical imaging diagnosis, accurate segmentation of the knee joint can help doctors better observe and diagnose lesions, thereby improving diagnostic accuracy and treatment effectiveness. Vision Mamba mainly relies on the State Space Model (SSM) for feature modeling, which excels at capturing global contextual information but cannot capture local texture features. Moreover, features of different scales are not effectively integrated, resulting in the model's weak segmentation ability on small-scale tissues (such as cartilage areas). To this end, this study proposed a novel multi-scale Vision Mamba Unet (VM-Unet) framework named MSPF-VM-Unet to perform the segmentation on the femur, tibia, femoral cartilage, and tibial cartilage in knee MRI images. The proposed MSPF-VM-Unet extends VM-Unet by introducing a designed multi-scale pyramid feature extraction network named MPSK, which synergizes multi-resolution feature extraction with channel-space attention. MPSK network enhances multi-scale local feature extraction through Selective Kernel (SK) convolution and pyramid pooling. The network merges the overall context information extracted by the Vision Mamba encoder to achieve the coordinated optimization of a multi-scale hierarchical feature fusion mechanism and global long-range dependency modeling. The results of the comparative experiments on the OAI-ZIB dataset indicate that MSPF-VM-Unet significantly improves the boundary accuracy and regional consistency of the MRI tibiofemoral joint tissue structure.

  • Research Article
  • Cite Count Icon 4
  • 10.1016/j.jag.2025.104731
OriMamba: Remote sensing oriented object detection with state space models
  • Sep 1, 2025
  • International Journal of Applied Earth Observation and Geoinformation
  • Zhanhao Xiao + 5 more

OriMamba: Remote sensing oriented object detection with state space models

  • Research Article
  • 10.1016/j.cmpb.2026.109515
DDVMM: A dual-branch pyramid model for mono-modal medical image registration.
  • Jun 15, 2026
  • Computer methods and programs in biomedicine
  • Maoyang Zou + 3 more

DDVMM: A dual-branch pyramid model for mono-modal medical image registration.

  • Research Article
  • Cite Count Icon 5
  • 10.1109/tim.2025.3533639
A State Space Model-Driven Multiscale Attention Method for Geological Hazard Segmentation
  • Jan 1, 2025
  • IEEE Transactions on Instrumentation and Measurement
  • Juan Yang + 3 more

High-precision segmentation of geological disasters plays a crucial role in disaster rescue, significantly contributing to improving rescue efficiency and optimizing the allocation of rescue resources. However, landslides and debris flows typically have irregular contours and arbitrary scopes, which cause most existing methods to suffer from poor performance. To address these issues, we propose a novel dual-path feature extraction architecture for geological hazard segmentation. First, the state space model-driven global multiscale attention module (SSMGMA) is used to model cross-scale long-range dependencies by powerful multiscale representation. Dilated convolutions are adopted to extract multiscale features, while the state space model (SSM) is incorporated to capture the global context and model cross-scale long-range dependencies. Consequently, the SSMGMA allows the proposed model to completely segment geological disasters. Subsequently, the high-frequency prompt encoder module (HFPE) is employed to alleviate the negative effects caused by irregular contour problems. The core idea of the HFPE is to effectively encode high-frequency information as the prompt to the decoder. Specifically, a well-designed encoding strategy is adopted in the HFPE, which can transform subtle variations in high-frequency information into precise locations of disaster areas. By combining the SSMGMA and HFPE, the proposed dual-path architecture leverages the advantages of multiscale features and high-frequency feature encoding, significantly improving the accuracy of disaster segmentation. Experimental results show that the proposed method has superior performance.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 11
  • 10.3390/app12052626
Multi-Scale Features for Transformer Model to Improve the Performance of Sound Event Detection
  • Mar 3, 2022
  • Applied Sciences
  • Soo-Jong Kim + 1 more

To alleviate the problem of performance degradation due to the varied sound durations of competing classes in sound event detection, we propose a method that utilizes multi-scale features for sound event detection. We employed a feature-pyramid component in a deep neural network architecture based on the Transformer encoder that is used to efficiently model the time correlation of sound signals because of its superiority over conventional recurrent neural networks, as demonstrated in recent studies. We used layers of convolutional neural networks to produce two-dimensional acoustic features that are input into the Transformer encoders. The outputs of the Transformer encoders at different levels of the network are combined to obtain the multi-scale features to feed the fully connected feed-forward neural network, which acts as the final classification layer. The proposed method is motivated by the idea that multi-scale features make the network more robust against the dynamic duration of the sound signals depending on their classes. We also applied the proposed method to a mean-teacher model, based on the Transformer encoder, to demonstrate its effectiveness on a large set of unlabeled data. We conducted experiments using the DCASE 2019 Task 4 dataset to evaluate the performance of the proposed method. The experimental results show that the proposed architecture outperforms the baseline network without multi-scale features.

  • Conference Article
  • Cite Count Icon 1
  • 10.2316/p.2010.675-037
Adaptive Predictive Control of Three - Tank - System
  • Jan 1, 2010
  • M Kubalčík + 1 more

This paper is focused in application of a self - tuning predictive controller for real - time control of a three - tank - system laboratory model. The objective laboratory model is a two input - two output (TITO) nonlinear system. It is based on experience with authentic industrial control applications. The controller integrates a predictive control synthesis based on a multivariable state - space model of the controlled system and an on - line identification of an ARX model corresponding to the state - space model. The model parameters are recursively estimated using the recursive least squares method with the directional forgetting. The control algorithm is based on the Generalised Predictive Control (GPC) method. The optimization was realized by minimization of a quadratic objective function. Results of real-time experiments are also included. approaches, the control actions are taken based on past errors. MPC uses also future values of the reference signals. The aim of this contribution is implementation of an adaptive predictive controller for control of the three - tank - system laboratory model. The design of the controller is based on a state - space model. An initial state - space model was constructed according to first principles and physical rules. The parameters of the system were not recognizable. Moreover, the laboratory model is a nonlinear system with variable parameters and its description by a linear model is valid only in a neighbourhood of a steady state. Self-tuning controllers (11), (12) are a possible approach to the control of this kind of system. However, the state - space description is not quite suitable for a recursive identification of the parameters of the process which is performed during control with self - tuning controllers. The state space model was then converted to a model in the form of difference equations. This model is suitable for the recursive identification. So the proposed approach combines both types of models. The state - space model is used for the controllers design and the corresponding input/output model for the estimation of the unknown parameters. Of course it is possible to base the controllers design on the input/output model as well. But the main theoretical results of predictive control come from a state space formulation, which can be used easily both for SISO and MIMO systems. It also enables to solve tasks which are unsolvable when using an input/output model. For example control with state constraints. Reverse conversion of the difference equations to the original state - space model is not possible. It is explained in section 3. An alternative state - space model was than established and used for the controllers design. This model corresponds to the original model despite the fact that it has a different structure. So it is possible to assume that this model describes main properties of the controlled process as well as the original model. The Generalised Predictive Control (GPC) method (13), (14) was then applied for the controllers design. In the optimization part of the algorithm a quadratic cost function was used. The algorithm takes into account constraints of manipulated variables. The recursive least squares method with the directional forgetting is used in the identification part.

  • Research Article
  • 10.3389/fpls.2026.1842426
Multi-scale feature fusion-based vision mamba for robust plant disease image classification on field-acquired plantdoc data
  • Jan 1, 2026
  • Frontiers in Plant Science
  • Shanjiang Zhang + 1 more

IntroductionExisting convolutional neural networks and Transformers cannot effectively capture fine-grained local lesion features and long-range contextual dependencies simultaneously in field-collected plant images. To address this research limitation, we aim to design an effective lightweight model suitable for plant disease identification in complex field scenarios.MethodsThis work proposes an improved Vision Mamba network for plant disease classification based on the challenging PlantDoc dataset. Three dedicated modules are embedded into the framework, including the Multi-Scale Feature Fusion Module (MFFM), Adaptive Channel Attention Mechanism (ACAM) and Lightweight Residual Connection (LRC). The MFFM fuses multi-scale texture, shape and semantic lesion features extracted from shallow, medium and deep network layers. The ACAM adaptively highlights disease-related feature channels and suppresses irrelevant background interference. The LRC structure is adopted to relieve the gradient vanishing problem existing in deep selective state space model (SSM) networks.ResultsExperimental results on the filtered PlantDoc dataset show that the presented model obtains an overall accuracy of 92.67%, macro precision of 91.83%, macro recall of 91.56% and macro F1-score of 91.70% on independent test samples, which outperforms the original Vision Mamba baseline by 5.33% in accuracy. Five-fold stratified cross-validation achieves stable accuracy at 92.41 ± 0.24%, and paired t-tests prove that the performance improvement is statistically significant with p<0.05. Ablation experiments confirm the combined contribution of the three designed modules.DiscussionError analysis and confusion matrix visualization reveal that the main classification errors are derived from high similarity among different plant disease categories. This study fully verifies the application potential of state space models in agricultural computer vision tasks. The proposed method can serve as an efficient technical scheme for intelligent identification of crop diseases and is well applicable to edge device deployment in precision agriculture practice.

  • Research Article
  • 10.3390/s25247414
HyMambaNet: Efficient Remote Sensing Water Extraction Method Combining State Space Modeling and Multi-Scale Features
  • Dec 5, 2025
  • Sensors (Basel, Switzerland)
  • Handan Liu + 6 more

Accurate segmentation of water bodies from high-resolution remote sensing imagery is crucial for water resource management and ecological monitoring. However, small and morphologically complex water bodies remain difficult to detect due to scale variations, blurred boundaries, and heterogeneous backgrounds. This study aims to develop a robust and scalable deep learning framework for high-precision water body extraction across diverse hydrological and ecological scenarios. To address these challenges, we propose HyMambaNet, a hybrid deep learning model that integrates convolutional local feature extraction with the Mamba state space model for efficient global context modeling. The network further incorporates multi-scale and frequency-domain enhancement as well as optimized skip connections to improve boundary precision and segmentation robustness. Experimental results demonstrate that HyMambaNet significantly outperforms existing CNN and Transformer-based methods. On the LoveHY dataset, it achieves 74.82% IoU and 88.87% F1-score, exceeding UNet by 7.49% IoU and 7.12% F1. On the LoveDA dataset, it attains 81.30% IoU and 89.99% F1-score, surpassing advanced models such as Deeplabv3+, AttenUNet, and TransUNet. These findings confirm that HyMambaNet provides an efficient and generalizable solution for large-scale water resource monitoring and ecological applications based on remote sensing imagery.

  • Research Article
  • Cite Count Icon 3
  • 10.1038/s41598-025-21837-2
Super Mamba feature enhancement framework for small object detection
  • Oct 23, 2025
  • Scientific Reports
  • Na Shi + 8 more

It is very challenging to accurately and timely detect small object containing dozens of pixels from infrared images. Compared with the complex background in infrared images taken by low-altitude drones, a framework is designed to learn a strong feature representation separating the object from the background, which usually leads to a large computational amount. In this paper, we proposed a Super Mamba (SMamba) framework for UAV infrared small object detection, which performs deep learning of nonlinear complex data. Our SMamba framework performs high resolution object detection on multi-scale objects, considering both detection accuracy and computational cost. First, the Receptive Field Attention Convolution (RFAConv) is used into the backbone network and replaced the commonly convolution, and the multi-scale features is adjusted through the dynamic receptive field to optimize the computing efficiency. Furthermore, the Spatial Attention Mechanism (SAM) and Squeeze-Excitation (SE) are added to the State Space Model (SSM) to achieve multi-scale and multi-feature extraction for small object. Moreover, in the neck, the Feature Enhancement Module (FEM) is introduced to Bidirectional Feature Pyramid Network (BiFPN) can enhance the local context information of small objects and improve the detection efficiency. The experimental results show that Super Mamba achieved more than 92% accuracy on VEDAI dataset (in terms of mAP@ 0.5), which is more than 20% higher than the existing large models such as Yolov5, Yolov8, and Yolov11. The pytorch code is available at: https://github.com/wolfololo/Super-Mamba-A-Framework-for-Small-Object-Detection-with-Enhanced-Detection.

  • Research Article
  • 10.1049/syb2.70044
MFS-Unet: A Multi-Path Vision Mamba Network for Precise Thyroid Nodule Segmentation.
  • Feb 1, 2026
  • IET systems biology
  • Shaoqiang Wang + 9 more

The automated segmentation of thyroid nodules from ultrasound images holds significant value in clinical diagnosis and treatment. However, achieving precise segmentation remains a substantial challenge due to issues such as blurred nodule boundaries, variable scales, image noise, and inaccurate annotations. To address these difficulties, this paper proposes a novel medical image segmentation network named MFS-Unet. The network introduces three innovative modules to enhance segmentation performance. First, we designed the multi-path vision mamba (MPV) module, which leverages the advantages of state space models (SSMs) to efficiently capture global contextual information and multi-scale features with linear computational complexity, effectively addressing the problem of significant variations in nodule size. Second, a feature gating (FG) module is deployed in the skip connections between the encoder and decoder. Through an attention mechanism, it dynamically screens and enhances features transmitted from the encoder, suppressing background noise and reinforcing key boundary information of the nodules. Finally, we propose a supervised label rectification (SLR) module, aimed at proactively handling the prevalent issue of label noise in training data. By dynamically adjusting loss weights during training, it guides the model to learn more robust feature representations. We conducted extensive experiments on three public thyroid ultrasound datasets: DDTI, TG3K, and TN3K. The results demonstrate that MFS-Unet achieves superior performance across all evaluation metrics compared with various state-of-the-art segmentation methods, proving its effectiveness and significant potential for precise thyroid nodule segmentation in complex ultrasound environments.

  • Research Article
  • Cite Count Icon 9
  • 10.1016/j.neunet.2025.107919
SCFMUNet: A fusion architecture based on multi-scale state space model and channel attention for medical image segmentation.
  • Dec 1, 2025
  • Neural networks : the official journal of the International Neural Network Society
  • Zhiyong Huang + 8 more

SCFMUNet: A fusion architecture based on multi-scale state space model and channel attention for medical image segmentation.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.4038/sljastats.v17i2.7872
State space versus SARIMA modeling of the Nigeria’s crude oil export
  • Nov 9, 2016
  • Sri Lankan Journal of Applied Statistics
  • Omorogbe Joseph Asemota

This paper analyzes the Nigeria’s crude oil export series using monthly data from January 1999 to December 2014. We employed the state space local level model with stochastic and deterministic seasonal to model the dynamic features in the Nigeria crude oil export. Our results clearly indicate that the local level model with deterministic seasonal is the most parsimonious model between the two state space models considered in this study. Also, a parsimonious SARIMA model is also fitted to the data. We compare the forecasting performance of the two parsimonious models and evaluate their forecasts using ex-post indicators such as mean absolute percentage error (MAPE), root mean square percentage error (RMSPE) and the Theil’s U statistic. The forecast analysis and evaluation results indicate that the state space local level model with deterministic seasonal outperforms the Box-Jenkins model in shorter and medium – range forecasting horizons. Howbeit, the forecast of the SARIMA model improves in the longer horizon. The Theil’s U statistic also indicates that the state space local level model with deterministic seasonal and SARIMA model outperform the naive model at most of the forecasting horizons. In conclusion, we recommend that the state space model with deterministic seasonal component should be used in shorter and medium range forecasting horizons of the Nigeria’s monthly crude oil export. Howbeit, for longer forecasting horizon, ten months and above, the seasonal ARIMA model should be considered.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant