Uneven success: automatic speech recognition and ethnicity-related dialects

Alicia Beckford Wassink,Cady Gansen,Isabel Bartholomew

doi:10.1016/j.specom.2022.03.009

Alicia Beckford Wassink, Cady Gansen + Show 1 more

Open Access

https://doi.org/10.1016/j.specom.2022.03.009

Copy DOI

Journal: Speech Communication	Publication Date: Mar 12, 2022
Citations: 14	License type: publisher-specific-oa

Affiliation: University of Washington, Seattle University

Abstract

Addressing racial bias in automatic speech recognition is an area of concern in fields associated with human-computer interaction. Research to date suggests that sociolinguistic variation, namely systematic sources of sociophonetic variation, has yet to be extensively exploited in acoustic model architectures. This paper reports a study that evaluates the performance of one ASR system for a multi-ethnic sample of speakers from the American Pacific Northwest (including Native American, African American, European American and ChicanX speakers). Using a sociophonetic approach to characterizing vocalic and consonantal variation, we ask which dialect features appear to be most challenging for our ASR system. We also ask which error types are particular to the four ethnic dialects sampled. Recordings of both conversational and read speech were coded for a common set of 18 sociophonetic variables with distinct phonetic profiles. Automatic transcription was achieved using CLOx, a custom-built ASR system created for sociolinguistic analysis. Normalized error frequency rates were compared across ethnic samples to evaluate CLOx performance. Nf error rates demonstrate clear differential performance in the ASR system, pointing to racial bias in system output. Specific predictions are made regarding approaches that might be taken to leverage sociophonetic knowledge to improve social dialect-recognition accuracy in ASR systems.

Full Text