Speaker Verification System Research Articles

The use of face masks has increased dramatically since the COVID-19 pandemic started in order to to curb the spread of the disease. Additionally, breakthrough infections caused by the Delta and Omicron variants have further increased the importance of wearing a face mask, even for vaccinated individuals. However, the use of face masks also induces attenuation in speech signals, and this change may impact speech processing technologies, e.g., automated speaker verification (ASV) and speech to text conversion. In this paper we examine Automatic Speaker Verification (ASV) systems against the speech samples in the presence of three different types of face mask: surgical, cloth, and filtered N95, and analyze the impact on acoustics and other factors. In addition, we explore the effect of different microphones, and distance from the microphone, and the impact of face masks when speakers use ASV systems in real-world scenarios. Our analysis shows a significant deterioration in performance when an ASV system encounters different face masks, microphones, and variable distance between the subject and microphone. To address this problem, this paper proposes a novel framework to overcome performance degradation in these scenarios by realigning the ASV system. The novelty of the proposed ASV framework is as follows: first, we propose a fused feature descriptor by concatenating the novel Ternary Deviated overlapping Patterns (TDoP), Mel Frequency Cepstral Coefficients (MFCC), and Gammatone Cepstral Coefficients (GTCC), which are used by both the ensemble learning-based ASV and anomaly detection system in the proposed ASV architecture. Second, this paper proposes an anomaly detection model for identifying vocal samples produced in the presence of face masks. Next, it presents a Peak Norm (PN) filter to approximate the signal of the speaker without a face mask in order to boost the accuracy of ASV systems. Finally, the features of filtered samples utilizing the PN filter and samples without face masks are passed to the proposed ASV to test for improved accuracy. The proposed ASV system achieved an accuracy of 0.99 and 0.92, respectively, on samples recorded without a face mask and with different face masks. Although the use of face masks affects the ASV system, the PN filtering solution overcomes this deficiency up to 4%. Similarly, when exposed to different microphones and distances, the PN approach enhanced system accuracy by up to 7% and 9%, respectively. The results demonstrate the effectiveness of the presented framework against an in-house prepared, diverse Multi Speaker Face Masks (MSFM) dataset, (IRB No. FY2021-83), consisting of samples of subjects taken with a variety of face masks and microphones, and from different distances.

Read full abstract

Background/Objectives: The anti-spoofing measures are blooming with an aim to protect the Automatic Speaker Verification systems from susceptible spoofing attacks. This review is an amalgam of the possible attack types, the datasets required, the renowned feature representation techniques, modeling algorithms involving machine learning, and score normalization techniques. Method/Findings: A detailed analysis of existing datasets is carried based on the total speaker samples, the number of speakers, and source of availability- open or licensed. This may foster choosing the right dataset for building the anti-spoofing frameworks. Further, the feature extraction schemes are elaborated with an intention to cover the vast span of features existing in various parts of raw speech for obtaining speaker-specific traits. Further, the machine learning algorithms ranging from discriminative to generative to mixed form are explored for seeking the right algorithm in specific attack conditions. On the whole, these analyses of existing features and machine learning algorithms together contribute to classifying the unknown test samples as genuine or spoofed. The score normalization techniques are also considered in this review to avoid any misclassifications and ultimately reduce the False Acceptance Ratios. The performance of any anti-spoofing speaker verification system may be evaluated using standard objective measures such are Equal Error Rate, False positive ratios, and graphical plots. These measures are briefly explained in this review. Overall, the critical analysis of individual methods-feature extraction, machine learning, score normalization, and all the anti-spoofing datasets are also discussed for giving a kick-start to any researcher beginning to explore in this direction. The shortcomings and risks involved in building an enhanced speaker verification system that is robust to almost all the attack types are listed in this article. The review of studies conducted so far has led to vital future directions that are enlisted in the concluding remarks of the article. Keywords: Automatic Speaker Verification; Spoofed Detection, AntiSpoofing, Voice Conversion, Speech Synthesis, Replay Speech

Read full abstract

Speaker Verification System Research Articles

Related Topics

Articles published on Speaker Verification System

Replay spoof detection for speaker verification system using magnitude-phase-instantaneous frequency and energy features

Toward Realigning Automatic Speaker Verification in the Era of COVID-19

Shouted and whispered speech compensation for speaker verification systems

A robust voice spoofing detection system using novel CLS-LBP features and LSTM

Attention-Based Temporal-Frequency Aggregation for Speaker Verification.

Voice spoofing detector: A unified anti-spoofing framework

Multi-task Learning-Based Spoofing-Robust Automatic Speaker Verification System

New replay attack detection using iterative adaptive inverse filtering and high frequency band

A hybrid noise robust model for multireplay attack detection in Automatic speaker verification systems

Generative and Discriminative Modelling of Linear Energy Sub-bands for Spoof Detection in Speaker Verification Systems

Uncertainty assessment for detection of spoofing attacks to speaker verification systems using a Bayesian approach

ADCF Loss Function for Deep Metric Learning in End-to-End Text-Dependent Speaker Verification Systems

A Survey on Text-Dependent and Text-Independent Speaker Verification

Optimizing Tandem Speaker Verification and Anti-Spoofing Systems

Verifying student identity in oral assessments with deep speaker

DeltaVLAD: An efficient optimization algorithm to discriminate speaker embedding for text-independent speaker verification

A Supervised Learning Method for Improving the Generalization of Speaker Verification Systems by Learning Metrics from a Mean Teacher

Static\u2013dynamic features and hybrid deep learning models based spoof detection system for ASV

A review on state-of-the-art Automatic Speaker verification system from spoofing and anti-spoofing perspective

Improving the potential of Enhanced Teager Energy Cepstral Coefficients (ETECC) for replay attack detection

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Speaker Verification System Research Articles

Related Topics

Articles published on Speaker Verification System

Replay spoof detection for speaker verification system using magnitude-phase-instantaneous frequency and energy features

Toward Realigning Automatic Speaker Verification in the Era of COVID-19

Shouted and whispered speech compensation for speaker verification systems

A robust voice spoofing detection system using novel CLS-LBP features and LSTM

Attention-Based Temporal-Frequency Aggregation for Speaker Verification.

Voice spoofing detector: A unified anti-spoofing framework

Multi-task Learning-Based Spoofing-Robust Automatic Speaker Verification System

New replay attack detection using iterative adaptive inverse filtering and high frequency band

A hybrid noise robust model for multireplay attack detection in Automatic speaker verification systems

Generative and Discriminative Modelling of Linear Energy Sub-bands for Spoof Detection in Speaker Verification Systems

Uncertainty assessment for detection of spoofing attacks to speaker verification systems using a Bayesian approach

ADCF Loss Function for Deep Metric Learning in End-to-End Text-Dependent Speaker Verification Systems

A Survey on Text-Dependent and Text-Independent Speaker Verification

Optimizing Tandem Speaker Verification and Anti-Spoofing Systems

Verifying student identity in oral assessments with deep speaker

DeltaVLAD: An efficient optimization algorithm to discriminate speaker embedding for text-independent speaker verification

A Supervised Learning Method for Improving the Generalization of Speaker Verification Systems by Learning Metrics from a Mean Teacher

Static\u2013dynamic features and hybrid deep learning models based spoof detection system for ASV

A review on state-of-the-art Automatic Speaker verification system from spoofing and anti-spoofing perspective

Improving the potential of Enhanced Teager Energy Cepstral Coefficients (ETECC) for replay attack detection