An extraction method of pop music singing beats based on audio features
An extraction method of pop music singing beats based on audio features
- Conference Article
- 10.1109/ccpr.2008.86
- Oct 1, 2008
Audio classification is based on audio features. The choice of audio features can reflect important audio classification features in time and frequency time. The extraction and analysis of audio features are the base and important of audio classification. The most important problem is to extract audio features effectively and make them mutual independence to reduce information redundancy. In this paper, combined with independent component analysis and rough set, a method for audio feature extraction is presented and it's proved better performance by experiments.
- Research Article
6
- 10.1007/s11517-006-0106-5
- Sep 26, 2006
- Medical & Biological Engineering & Computing
Science of human identification using physiological characteristics or biometry has been of great concern in security systems. However, robust multimodal identification systems based on audio-visual information has not been thoroughly investigated yet. Therefore, the aim of this work to propose a model-based feature extraction method which employs physiological characteristics of facial muscles producing lip movements. This approach adopts the intrinsic properties of muscles such as viscosity, elasticity, and mass which are extracted from the dynamic lip model. These parameters are exclusively dependent on the neuro-muscular properties of speaker; consequently, imitation of valid speakers could be reduced to a large extent. These parameters are applied to a hidden Markov model (HMM) audio-visual identification system. In this work, a combination of audio and video features has been employed by adopting a multistream pseudo-synchronized HMM training method. Noise robust audio features such as Mel-frequency cepstral coefficients (MFCC), spectral subtraction (SS), and relative spectra perceptual linear prediction (J-RASTA-PLP) have been used to evaluate the performance of the multimodal system once efficient audio feature extraction methods have been utilized. The superior performance of the proposed system is demonstrated on a large multispeaker database of continuously spoken digits, along with a sentence that is phonetically rich. To evaluate the robustness of algorithms, some experiments were performed on genetically identical twins. Furthermore, changes in speaker voice were simulated with drug inhalation tests. In 3 dB signal to noise ratio (SNR), the dynamic muscle model improved the identification rate of the audio-visual system from 91 to 98%. Results on identical twins revealed that there was an apparent improvement on the performance for the dynamic muscle model-based system, in which the identification rate of the audio-visual system was enhanced from 87 to 96%.
- Research Article
3
- 10.5573/ieiespc.2019.8.2.100
- Apr 30, 2019
- IEIE Transactions on Smart Processing & Computing
Recently, there has been increasing interest in artificial intelligence and machine learning, where sentiment analysis has received considerable attention. In several studies, emotional states have been recognized using audio, text, or bio-signals that induce emotions, with audio being the most typical. There are several audio features, such as rhythm, dynamics, melody, harmony, and tonal color. The aim of our paper is finding critical audio features for effective emotion recognition. To do this, we select the existing audio features from elements of music, and investigate critical features using an iterative feature extraction method. For objective evaluation, the International Affective Digital Sounds system was used for training and testing. Crossvalidation evaluated the method in terms of classifier accuracy and computational complexity, and the results indicate the critical features for emotion classification.
- Conference Article
8
- 10.1109/acssc.2012.6489302
- Nov 1, 2012
This paper aims at comparing the discrimination between audio, 2D-based visual and 3D-based visual features for the speech recognition purpose. The audio and visual feature extraction schemes and several feature selection techniques are described first in this paper. With the application of the described feature extraction and selection methods, several experiments are conducted to compare the discrimination of the audio features, the 2D visual features and the 3D visual features for the hVd words classification task. In our study, it is found that the 3D visual features have more separability than the 2D visual features, so that the 3D-based audio-visual speech recognition may achieve more desirable results than the traditional 2D-based counterpart.
- Conference Article
- 10.1109/icee.2018.8472620
- May 1, 2018
From searching music with smartphones to broadcast monitoring by radio channels, audio identification systems are being used more in recent years. Design of such systems may differ when the problem domain changes, since each environment has special conflicting constraints to consider, like required speed and robustness to signal degradations. In this paper, a widely used audio identification system originally developed by Haitsma and Kalker is analyzed from a signal processing point of view and the fingerprint (audio feature) extraction method is modified. By adding a flexible filter to the fingerprint extraction method, the original system can be tuned to work in different domains. In order to optimize the filters for each domain, Pareto optimization is used with the multi-obj ective genetic algorithm. The novel approach to make an adaptive audio identification system leads to the fact that instead of making fundamentally different systems for different domains, one general system can be developed and then modified to meet each domain-specific constraints. By using optimized filters, the average accuracy was increased to 94.28% from 69.78% which is the average accuracy of the original system.
- Conference Article
12
- 10.1117/12.586533
- Jan 17, 2005
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
In this paper, we present an automatic extraction of goal events in soccer videos by using audio track features alone without relying on expensive-to-compute video track features. The extracted goal events can be used for high-level indexing and selective browsing of soccer videos. The detection of soccer video highlights using audio contents comprises three steps: 1) extraction of audio features from a video sequence, 2) event candidate detection of highlight events based on the information provided by the feature extraction Methods and the Hidden Markov Model (HMM), 3) goal event selection to finally determine the video intervals to be included in the summary. For this purpose we compared the performance of the well known Mel-scale Frequency Cepstral Coefficients (MFCC) feature extraction method vs. MPEG-7 Audio Spectrum Projection feature (ASP) extraction method based on three different decomposition methods namely Principal Component Analysis( PCA), Independent Component Analysis (ICA) and Non-Negative Matrix Factorization (NMF). To evaluate our system we collected five soccer game videos from various sources. In total we have seven hours of soccer games consisting of eight gigabytes of data. One of five soccer games is used as the training data (e.g., announcers' excited speech, audience ambient speech noise, audience clapping, environmental sounds). Our goal event detection results are encouraging.
- Conference Article
- 10.1145/3665348.3665373
- May 10, 2024
Although the subjective evaluation method is the most accurate evaluation method, how to use the advantages of the subjective evaluation method to study a practical research method that can meet the subjective perception of the audio mapping library has become an urgent problem. Fast and efficient audio signal classification, feature recognition and signal reconstruction have become a common problem in audio signal processing. This paper combs, analyzes and summarizes audio signal classification technology, audio signal feature extraction and audio signal reconstruction methods, and analyzes the recognition methods and recognition effects based on the research status and achievements of domestic audio signal processing technology in the past decade. It is hoped that the research of this paper can provide reference for the theoretical research and engineering application of audio mapping library.
- Research Article
- 10.1002/pra2.1040
- Oct 1, 2024
- Proceedings of the Association for Information Science and Technology
ABSTRACTUp to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio modalities leads to deficiencies in information richness and overlooks potential correlations, thereby constraining the model's ability to learn representations of the data and the accuracy of model predictions. Furthermore, the currently available multimodal datasets for keyword extraction task are particularly scarce, further hindering the progress of research on multimodal keyword extraction task. Therefore, this study constructs a multimodal dataset of academic paper consisting of 1,000 samples, with each sample containing paper text, images, audios and keywords. Based on unsupervised and supervised methods of keyword extraction, experiments are conducted using textual data from papers, as well as text extracted from images and audio. The aim is to investigate the differences in performance in keyword extraction task with respect to different modal information and the fusion of multimodal information. The experimental results indicate that text from different modalities exhibits distinct characteristics in the model. The concatenation of paper text, image text and audio text can effectively enhance the keyword extraction performance of academic papers.
- Conference Article
4
- 10.1109/ism.2014.81
- Dec 1, 2014
Automatic recognition of human emotional states is an important task for efficient human-machine communication. Most of existing works focus on the recognition of emotional states using audio signals alone, visual signals alone, or both. Here we propose empirical methods for feature extraction and classifier optimization that consider the temporal aspects of audio signals and introduce our framework to efficiently recognize human emotional states from audio signals. The framework is based on the prediction of input audio clips that are described using representative low-level features. In the experiments, seven (7) discrete emotional states (anger, fear, boredom, disgust, happiness, sadness, and neutral) from EmoDB dataset, are recognized and tested based on nineteen (19) audio features (15 standalone, 4 joint) by using the Support Vector Machine (SVM) classifier. Extensive experiments have been conducted to demonstrate the effect of feature extraction and classifier optimization methods to the recognition accuracy of the emotional states. Our experiments show that, feature extraction and classifier optimization procedures lead to significant improvement of over 11% in emotion recognition. As a result, the overall recognition accuracy achieved for seven emotions in the EmoDB dataset is 83.33% compared to the baseline accuracy of 72.22%.
- Research Article
22
- 10.3390/s22186966
- Sep 14, 2022
- Sensors
Aphasia is a type of speech disorder that can cause speech defects in a person. Identifying the severity level of the aphasia patient is critical for the rehabilitation process. In this research, we identify ten aphasia severity levels motivated by specific speech therapies based on the presence or absence of identified characteristics in aphasic speech in order to give more specific treatment to the patient. In the aphasia severity level classification process, we experiment on different speech feature extraction techniques, lengths of input audio samples, and machine learning classifiers toward classification performance. Aphasic speech is required to be sensed by an audio sensor and then recorded and divided into audio frames and passed through an audio feature extractor before feeding into the machine learning classifier. According to the results, the mel frequency cepstral coefficient (MFCC) is the most suitable audio feature extraction method for the aphasic speech level classification process, as it outperformed the classification performance of all mel-spectrogram, chroma, and zero crossing rates by a large margin. Furthermore, the classification performance is higher when 20 s audio samples are used compared with 10 s chunks, even though the performance gap is narrow. Finally, the deep neural network approach resulted in the best classification performance, which was slightly better than both K-nearest neighbor (KNN) and random forest classifiers, and it was significantly better than decision tree algorithms. Therefore, the study shows that aphasia level classification can be completed with accuracy, precision, recall, and F1-score values of 0.99 using MFCC for 20 s audio samples using the deep neural network approach in order to recommend corresponding speech therapy for the identified level. A web application was developed for English-speaking aphasia patients to self-diagnose the severity level and engage in speech therapies.
- Conference Article
- 10.1109/apccas55924.2022.10090286
- Nov 11, 2022
For keyword spotting (KWS) systems that usually work in mobile devices, a low-complexity design is essential for long stand-by time. Audio feature extraction and classifier modeling are the two main components of KWS systems. Log Mel-Frequency Spectral Coefficient (MFSC) is common for audio feature extraction due to its low complexity and good performance. Binary neural network (BNN) classifier, which owns binary weights and activations and performs convolution with XNOR, is applicable to low-complexity KWS applications. However, audio features are usually quantized with multiple-bit binary code to maintain high classification accuracy, which requires addition (ADD) operations in the first convolutional layer of the BNN model. Therefore, both XNOR and ADD units are needed in the BNN accelerator. To further reduce the complexity of KWS systems, we propose a new feature extraction method: Thermometer Codes of MFSC (MFSC-TC). Without LOG and DELTA operations, it is simpler than other MFSC-based methods. More importantly, convolution of all layers can be done by XNOR units due to the feature of thermometer code. The experiments with the Google Speech Commands dataset validate that the MFSC-TC-based BNN models outperform the models with more layers using other feature extraction methods.
- Conference Article
16
- 10.1109/icassp.2011.5946809
- May 1, 2011
We describe and evaluate our toolkit openBliSSART (open-source Blind Source Separation for Audio Recognition Tasks), which is the C++ framework and toolbox that we have successfully used in a multiplicity of research on blind audio source separation and feature extraction. To our knowledge, it provides the first open-source implementation of a widely applicable algorithmic framework based on non-negative matrix factorization (NMF), including several preprocessing, factorization, and signal reconstruction algorithms for monaural signals. Apart from blind source separation using supervised and unsupervised NMF, we show how the framework is useful for the increasingly popular audio feature extraction methods by NMF. Furthermore, we point out a numerical optimization for NMF, and show that NMF source separation in real-time on a desktop PC is feasible with our implementation. We conclude with an evaluation of our toolkit on supervised speaker separation, demonstrating how our algorithmic framework allows to tune the real-time factors to the desired perceptual quality.
- Research Article
- 10.52783/cana.v32.4274
- Mar 12, 2025
- Communications on Applied Nonlinear Analysis
Emotion recognition based on multimodal data (e.g., video, audio, text, etc.) is a highly demanding and significant research field with numerous applications. This research rigorously explores model level fusion to find the best multifunctional model combining audio and visual modalities for emotion identification. Specifically, it proposes novel feature extractor networks for both audio and video data. This research presents a comprehensive approach to multimodal emotion recognition, utilizing state-of-the-art feature extraction methods tailored to each modality. For text data, we implement the Assimilated N-gram Approach (ANA) to effectively capture contextual information. Audio features are extracted using Mel-Frequency Cepstral Coefficients (MFCC), ideal for capturing spectral characteristics in speech. Visual features are derived using Squeezenet, a deep learning architecture optimized for efficient and informative visual data representation. To integrate the extracted features from text, audio, and visual modalities, propose a multimodal data fusion strategy that combines information across modalities, thereby enhancing the overall representation of emotional cues. In the classification stage, employ Capsule Net, a novel neural network architecture adept at capturing hierarchical relationships and spatial hierarchies within data, making it well-suited for handling complex multimodal data. To further optimize the performance of the Capsule Net classifier, utilize hyper parameter tuning through the Sand Cat Swarm Optimization (SCSO) algorithm. SCSO, a metaheuristic optimization technique inspired by the behavior of sand cats, iteratively updates candidate solutions to converge towards optimal hyperparameter configurations. Using the Multimodal Emotion Lines Dataset (MELD), our approach achieved an accuracy of 98.91%, precision of 98.83%, recall of 99.04%, and F-measure of 98.94. These results highlight the effectiveness of our multimodal framework in emotion recognition tasks.
- Research Article
1
- 10.1155/2022/6713468
- Apr 11, 2022
- Wireless Communications and Mobile Computing
The piano is known as the king of musical instruments for its rich expressiveness. Pianists get rich emotion with tone control. Touch key mode is the primary method of tone control. The purpose of this paper is to conduct research and analysis on different keying modes and brain-sound image establishment modes in piano performance. In this paper, we first propose a fuzzy mathematical proximity comparison method and an audio feature extraction method. Most of the audio features are derived from speech recognition tasks. They can reduce the original waveform sampling signal, thereby accelerating the machine’s understanding of the semantic meaning of audio. Using comparative analysis method, fuzzy mathematical proximity comparison method, and audio feature extraction as research methods, a model analysis research model based on different key touch methods in piano playing was established, and the relationship between piano key touch and brain sound image was studied. The experimental results of this paper show that after using the research method in this paper, the error rate of the data is controlled within 5%. Compared with previous research methods, the error rate is lower and has certain practical value.
- Research Article
17
- 10.1121/1.4978245
- Mar 1, 2017
- The Journal of the Acoustical Society of America
By varying the dynamics in a musical performance, the musician can convey structure and different expressions. Spectral properties of most musical instruments change in a complex way with the performed dynamics, but dedicated audio features for modeling the parameter are lacking. In this study, feature extraction methods were developed to capture relevant attributes related to spectral characteristics and spectral fluctuations, the latter through a sectional spectral flux. Previously, ground truths ratings of performed dynamics had been collected by asking listeners to rate how soft/loud the musicians played in a set of audio files. The ratings, averaged over subjects, were used to train three different machine learning models, using the audio features developed for the study as input. The highest result was produced from an ensemble of multilayer perceptrons with an R2 of 0.84. This result seems to be close to the upper bound, given the estimated uncertainty of the ground truth data. The result is well above that of individual human listeners of the previous listening experiment, and on par with the performance achieved from the average rating of six listeners. Features were analyzed with a factorial design, which highlighted the importance of source separation in the feature extraction.