Coupling a Generative Model With a Discriminative Learning Framework for Speaker Verification

Xugang Lu,Peng Shen,Yu Tsao,Hisashi Kawai

doi:10.1109/taslp.2021.3129360

Abstract

The task of speaker verification (SV) is to decide whether an utterance is spoken by a target or an imposter speaker. In most studies of SV, a log-likelihood ratio (LLR) score is estimated based on a generative probability model on speaker features, and compared with a threshold for making a decision. However, the generative model usually focuses on individual feature distributions, does not have the discriminative feature selection ability, and is easy to be distracted by nuisance features. The SV, as a hypothesis test, could be formulated as a binary discrimination task where neural network based discriminative learning could be applied. In discriminative learning, the nuisance features could be removed with the help of label supervision. However, discriminative learning pays more attention to classification boundaries, and is prone to overfitting to a training set which may result in bad generalization on a test set. In this paper, we propose a hybrid learning framework, i.e., coupling a joint Bayesian (JB) generative model structure and parameters with a neural discriminative learning framework for SV. In the hybrid framework, a two-branch Siamese neural network is built with dense layers that are coupled with factorized affine transforms as used in the JB model. The LLR score estimation in the JB model is formulated according to the distance metric in the discriminative learning framework. By initializing the two-branch neural network with the generatively learned model parameters of the JB model, we further train the model parameters with the pairwise samples as a binary discrimination task. Moreover, a direct evaluation metric (DEM) in SV based on minimum empirical Bayes risk (EBR) is designed and integrated as an objective function in the discriminative learning. We carried out SV experiments on Speakers in the wild (SITW) and Voxceleb. Experimental results showed that our proposed model improved the performance with a large margin compared with state of the art models for SV.

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE/ACM transactions on audio, speech, and language processing	Publication Date: Jan 1, 2021
Citations: 1	License type: publisher-specific, author manuscript

R Discovery Prime

R Discovery Prime

Coupling a Generative Model With a Discriminative Learning Framework for Speaker Verification

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM transactions on audio, speech, and language processing

Lead the way for us

Similar Papers

A new SVM approach to speaker identification and verification using probabilistic distance kernels
Pedro J Moreno ... Purdy P Ho
-
Pedro J Moreno, et. al.Pedro J Moreno ... Purdy P Ho
01 Sep 2003
01 Sep 2003

A Double Joint Bayesian Approach for J-Vector Based Text-dependent Speaker Verification
Ziqiang Shi ... Mengjiao Wang
-
Ziqiang Shi, et. al.Ziqiang Shi ... Mengjiao Wang
26 Jun 2018
26 Jun 2018

Semi-supervised generative and discriminative adversarial learning for motor imagery-based brain\u2013computer interface
Wonjun Ko ... Jee Seok Yoon
Scientific Reports | VOL. 12
Wonjun Ko, et. al.Wonjun Ko ... Jee Seok Yoon
17 Mar 2022
Semi-supervised generative and discriminative adversarial learning for motor imagery-based brain\u2013computer interface
Wonjun Ko ... Jee Seok Yoon

Investigating loss functions and optimization methods for discriminative learning of label sequences
Yasemin Altun ... Thomas Hofmann
-
Yasemin Altun, et. al.Yasemin Altun ... Thomas Hofmann
01 Jan 2003
01 Jan 2003

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Coupling a Generative Model With a Discriminative Learning Framework for Speaker Verification

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM transactions on audio, speech, and language processing