H-VECTORS: Improving the robustness in utterance-level speaker embeddings using a hierarchical attention model

Qiang Huang,Thomas Hain,Yanpei Shi

doi:10.1016/j.neunet.2021.05.024

Qiang Huang, Thomas Hain + Show 1 more

Open Access

https://doi.org/10.1016/j.neunet.2021.05.024

Copy DOI

Journal: Neural Networks	Publication Date: May 25, 2021
Citations: 9	License type: public-domain

Affiliation: University of Sheffield

Abstract

In this paper, a hierarchical attention network is proposed to generate robust utterance-level embeddings (H-vectors) for speaker identification and verification. Since different parts of an utterance may have different contributions to speaker identities, the use of hierarchical structure aims to learn speaker related information locally and globally. In the proposed approach, frame-level encoder and attention are applied on segments of an input utterance and generate individual segment vectors. Then, segment level attention is applied on the segment vectors to construct an utterance representation. To evaluate the quality of the learned utterance-level speaker embeddings on speaker identification and verification, the proposed approach is tested on several benchmark datasets, such as the NIST SRE2008 Part1, the Switchboard Cellular (Part1), the CallHome American English Speech ,the Voxceleb1 and Voxceleb2 datasets. In comparison with some strong baselines, the obtained results show that the use of H-vectors can achieve better identification and verification performances in various acoustic conditions.

Full Text