Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition

Mengzhe Geng,Shujie Hu,Helen Meng,Tianzi Wang,Xunying Liu,Zi Ye,Guinan Li,Xurong Xie

doi:10.1109/taslp.2022.3195113

Abstract

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. Sources of heterogeneity commonly found in normal speech including accent or gender, when further compounded with the variability over age and speech pathology severity level, create large diversity among speakers. To this end, speaker adaptation techniques play a key role in personalization of ASR systems for such users. Motivated by the spectro-temporal level differences between dysarthric, elderly and normal speech that systematically manifest in articulatory imprecision, decreased volume and clarity, slower speaking rates and increased dysfluencies, novel spectro-temporal subspace basis deep embedding features derived using SVD speech spectrum decomposition are proposed in this paper to facilitate auxiliary feature based speaker adaptation of state-of-the-art hybrid DNN/TDNN and end-to-end Conformer speech recognition systems. Experiments were conducted on four tasks: the English UASpeech and TORGO dysarthric speech corpora; the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The proposed spectro-temporal deep feature adapted systems outperformed baseline i-Vector and x-Vector adaptation by up to 2.63% absolute (8.63% relative) reduction in word error rate (WER). Consistent performance improvements were retained after model based speaker adaptation using learning hidden unit contributions (LHUC) was further applied. The best speaker adapted system using the proposed spectral basis embedding features produced the lowest published WER of 25.05% on the UASpeech test set of 16 dysarthric speakers.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM Transactions on Audio, Speech, and Language Processing

Lead the way for us

Journal: IEEE/ACM Transactions on Audio, Speech, and Language Processing	Publication Date: Jan 1, 2022
Citations: 9

Similar Papers

Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition
Mengzhe Geng ... Zengrui Jin
-
Mengzhe Geng, et. al.Mengzhe Geng ... Zengrui Jin
30 Aug 2021
30 Aug 2021

Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
Xurong Xie ... Lan Wang
-
Xurong Xie, et. al.Xurong Xie ... Lan Wang
30 Aug 2021
30 Aug 2021

Recent Progress in the CUHK Dysarthric Speech Recognition System
Shansong Liu ... Mingyu Cui
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 29
Shansong Liu, et. al.Shansong Liu ... Mingyu Cui
01 Jan 2020
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 29

Auxiliary Networks for Joint Speaker Adaptation and Speaker Change Detection
Leda Sari ... Samuel Thomas
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 29
Leda Sari, et. al.Leda Sari ... Samuel Thomas
25 Nov 2020
IEEE/ACM Transactions on Audio, Speech, and Language Processing | VOL. 29

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech Recognition

Abstract

Talk to us

Similar Papers

More From: IEEE/ACM Transactions on Audio, Speech, and Language Processing