A Voice User Interface on the Edge for People with Speech Impairments

Davide Mulfari,Massimo Villari

doi:10.3390/electronics13071389

Abstract

Nowadays, fine-tuning has emerged as a powerful technique in machine learning, enabling models to adapt to a specific domain by leveraging pre-trained knowledge. One such application domain is automatic speech recognition (ASR), where fine-tuning plays a crucial role in addressing data scarcity, especially for languages with limited resources. In this study, we applied fine-tuning in the context of atypical speech recognition, focusing on Italian speakers with speech impairments, e.g., dysarthria. Our objective was to build a speaker-dependent voice user interface (VUI) tailored to their unique needs. To achieve this, we harnessed a pre-trained OpenAI’s Whisper model, which has been exposed to vast amounts of general speech data. However, to adapt it specifically for disordered speech, we fine-tuned it using our private corpus including 65 K voice recordings contributed by 208 speech-impaired individuals globally. We exploited three variants of the Whisper model (small, base, tiny), and by evaluating their relative performance, we aimed to identify the most accurate configuration for handling disordered speech patterns. Furthermore, our study dealt with the local deployment of the trained models on edge computing nodes, with the aim to realize custom VUIs for persons with impaired speech.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Voice User Interface on the Edge for People with Speech Impairments

Abstract

Talk to us

Similar Papers

More From: Electronics

Lead the way for us

Journal: Electronics	Publication Date: Apr 7, 2024
License type: CC BY 4.0

Similar Papers

Severity-Based Adaptation with Limited Data for ASR to Aid Dysarthric Speakers
Mumtaz Begum Mustafa ... Chng Eng Siong
PLoS ONE | VOL. 9
Mumtaz Begum Mustafa, et. al.Mumtaz Begum Mustafa ... Chng Eng Siong
23 Jan 2014
PLoS ONE | VOL. 9

Deep Learning in Intelligent Power and Energy Systems
Bruno Mota ... Tiago Pinto
-
Bruno Mota, et. al.Bruno Mota ... Tiago Pinto
02 Dec 2022
02 Dec 2022

Review of Machine and Deep Learning Techniques in Epileptic Seizure Detection using Physiological Signals and Sentiment Analysis
Deba Prasad Dash ... Mohammad R Khosravi
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 23
Deba Prasad Dash, et. al.Deba Prasad Dash ... Mohammad R Khosravi
15 Jan 2024
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 23

Validity of Off-the-Shelf Automatic Speech Recognition for Assessing Speech Intelligibility and Speech Severity in Speakers With Amyotrophic Lateral Sclerosis.
Sarah E Gutz ... Kaila L Stipancic
Journal of Speech, Language, and Hearing Research | VOL. 65
Sarah E Gutz, et. al.Sarah E Gutz ... Kaila L Stipancic
27 May 2022
Journal of Speech, Language, and Hearing Research | VOL. 65

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Voice User Interface on the Edge for People with Speech Impairments

Abstract

Talk to us

Similar Papers

More From: Electronics