Speaker Classification Research Articles

This study examines the accuracy of Interaction Detection in Early Childhood Settings (IDEAS), a program that automatically transcribes audio files and estimates linguistic units relevant to speech-language therapy, including part-of-speech units that represent features of language complexity, such as adjectives and coordinating conjunctions. Forty-five video-recorded speech-language therapy sessions involving 27 speech-language pathologists (SLPs) and 56 children were used. The F measure determines the accuracy of IDEAS diarization (i.e., speech segmentation and speaker classification). Two additional evaluation metrics, namely, median absolute relative error and correlation, indicate the accuracy of IDEAS for the estimation of linguistic units as compared with two conditions, namely, Oracle (manual diarization) and Voice Type Classifier (existing diarizer with acceptable accuracy). The high F measure for SLP talk data suggests high accuracy of IDEAS diarization for SLP talk but less so for child talk. These differences are reflected in the accuracy of IDEAS linguistic unit estimates. IDEAS median absolute relative error and correlation values for nine of the 10 SLP linguistic unit estimates meet the accuracy criteria, but none of the child linguistic unit estimates meet these criteria. The type of linguistic units also affects IDEAS accuracy. IDEAS was tailored to educational settings to automatically convert audio recordings into text and to provide linguistic unit estimates in speech-language therapy sessions and classroom settings. Although not perfect, IDEAS is reliable in automatically capturing and returning linguistic units, especially in SLP talk, that are relevant in research and practice. The tool offers a way to automatically measure SLP talk in clinical settings, which will support research seeking to understand how SLP talk influences children's language growth.

Read full abstract

AbstractHuman speech signals may contain specific information regarding a speaker's characteristics, and these signals can be very useful in applications involving interactive voice response (IVR) and automatic speech recognition (ASR). For IVR and ASR applications, speaker classification into different ages and gender groups can be applied in human–machine interaction or computer‐based interaction systems for customised advertisement, translation (text generation), machine dialog systems, or self‐service applications. Hence, an IVR‐based system dictates that ASR should function through users' voices (specific voice‐frequency bands) to identify customers' age and gender and interact with a host system. In the present study, we intended to combine a pitch detection (PD)‐based extractor and a voice classifier for gender identification. The Yet Another Algorithm for Pitch Tracking (YAAPT)‐based PD method was designed to extract the voice fundamental frequency (F0) from non‐stationary speaker's voice signals, allowing us to achieve gender identification, by distinguishing differences in F0 between adult females and males, and classify voices into adult and children groups. Then, in vowel voice signal classification, a one‐dimensional (1D) convolutional neural network (CNN), consisted of a multi‐round 1D kernel convolutional layer, a 1D pooling process, and a vowel classifier that could preliminary divide feature patterns into three level ranges of F0, including adult and children groups. Consequently, a classifier was used in the classification layer to identify the speakers' gender. The proposed PD‐based extractor and voice classifier could reduce complexity and improve classification efficiency. Acoustic datasets were selected from the Hillenbrand database for experimental tests on 12 vowels classifications, and K‐fold cross‐validations were performed. The experimental results demonstrated that our approach is a very promising method to quantify the proposed classifier's performance in terms of recall (%), precision (%), accuracy (%), and F1 score.

Read full abstract

Speaker Classification Research Articles

Related Topics

Articles published on Speaker Classification

Decoding Silent Speech Cues From Muscular Biopotential Signals for Efficient Human‐Robot Collaborations

An Automated PSO-CFNN Approach for Speaker Classification and Identification Using a Deep Learning Technique

Accuracy of Automatic Processing of Speech-Language Pathologist and Child Talk During School-Based Therapy Sessions.

Whisper-SV: Adapting Whisper for low-data-resource speaker verification

Effects of the landline telephone filter and linguistic context on speaker-dependent variability in /s/

Investigating children's interactions in preschool classrooms: An overview of research using automated sensing technologies

Vowel classification with combining pitch detection and one‐dimensional convolutional neural network based classifier for gender identification

Overview of Voice Conversion Methods Based on Deep Learning

SAPBERT: Speaker-Aware Pretrained BERT for Emotion Recognition in Conversation

A multi-tasking model of speaker-keyword classification for keeping human in the loop of drone-assisted inspection

Novel voiceprint using ensembled Mel-Chromagram for speaker recognition

Multi-Task Semi-Supervised Adversarial Autoencoding for Speech Emotion Recognition

The role of perceived ethnicity in speech processing: Insights from diverse populations and methods

S-Vectors and TESA: Speaker Embeddings and a Speaker Authenticator Based on Transformer Encoder

Disentangling Style and Speaker Attributes for TTS Style Transfer

Do faces speak volumes? Social expectations in speech comprehension and evaluation across three age groups.

Acoustic and speaker variation in Dutch /n/ and /m/ as a function of phonetic context and syllabic position

Development of a regional voice dataset and speaker classification based on machine learning

DNN and i-vector combined method for speaker recognition on multi-variability environments

An Accent Marking Algorithm of English Conversion System Based on Morphological Rules

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Speaker Classification Research Articles

Related Topics

Articles published on Speaker Classification

Decoding Silent Speech Cues From Muscular Biopotential Signals for Efficient Human‐Robot Collaborations

An Automated PSO-CFNN Approach for Speaker Classification and Identification Using a Deep Learning Technique

Accuracy of Automatic Processing of Speech-Language Pathologist and Child Talk During School-Based Therapy Sessions.

Whisper-SV: Adapting Whisper for low-data-resource speaker verification

Effects of the landline telephone filter and linguistic context on speaker-dependent variability in /s/

Investigating children's interactions in preschool classrooms: An overview of research using automated sensing technologies

Vowel classification with combining pitch detection and one‐dimensional convolutional neural network based classifier for gender identification

Overview of Voice Conversion Methods Based on Deep Learning

SAPBERT: Speaker-Aware Pretrained BERT for Emotion Recognition in Conversation

A multi-tasking model of speaker-keyword classification for keeping human in the loop of drone-assisted inspection

Novel voiceprint using ensembled Mel-Chromagram for speaker recognition

Multi-Task Semi-Supervised Adversarial Autoencoding for Speech Emotion Recognition

The role of perceived ethnicity in speech processing: Insights from diverse populations and methods

S-Vectors and TESA: Speaker Embeddings and a Speaker Authenticator Based on Transformer Encoder

Disentangling Style and Speaker Attributes for TTS Style Transfer

Do faces speak volumes? Social expectations in speech comprehension and evaluation across three age groups.

Acoustic and speaker variation in Dutch /n/ and /m/ as a function of phonetic context and syllabic position

Development of a regional voice dataset and speaker classification based on machine learning

DNN and i-vector combined method for speaker recognition on multi-variability environments

An Accent Marking Algorithm of English Conversion System Based on Morphological Rules