A Corpus of Regional American Language from YouTube

Steven Coats

doi:10.5617/dhnbpub.11083

Abstract

Recent years have seen an increase in the number of corpora of regional language variation for English, allowing new types of aggregate analysis to be conducted. While the creation of a corpus from written language material is relatively straightforward, transcribing speech is time-consuming, and thus there are no large corpora of transcribed American speech with broad geographic coverage. This paper describes the creation of a new corpus of regional American English from the automatically generated captions of videos from YouTube channels with a local American focus – mainly channels of regional and local government entities or civic organizations. The corpus, which consists of transcripts of over 29,267 hours of spoken language, will enable the analysis of regional patterns of lexical, morphosyntactic, and other types of variation in spoken American English. Exploratory analysis and mapping of the corpus data indicates regional variation in spoken language is evident.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Corpus of Regional American Language from YouTube

Abstract

Talk to us

Similar Papers

More From: Digital Humanities in the Nordic and Baltic Countries Publications

Lead the way for us

Journal: Digital Humanities in the Nordic and Baltic Countries Publications	Publication Date: May 17, 2019
License type: CC BY 4.0

Similar Papers

Sociophonetic analysis of vowel variation in African American English in the Southern United States
Yolanda F Holt
-
Yolanda F HoltYolanda F Holt
01 Jan 2015
01 Jan 2015

American and Irish English speakers’ perceptions of the final particles so and but
Mitsuko Narita Izutsu ... Katsunobu Izutsu
World Englishes | VOL. 41
Mitsuko Narita Izutsu, et. al.Mitsuko Narita Izutsu ... Katsunobu Izutsu
17 Sep 2020
World Englishes | VOL. 41

The discrimination and categorization of German /r/ allophones by American English speakers.
Dilara Tepeli
The Journal of the Acoustical Society of America | VOL. 129
Dilara TepeliDilara Tepeli
01 Apr 2011
The Journal of the Acoustical Society of America | VOL. 129

Expressing smells in (American) English
Doris Eveline Schönefeld
Corpus Linguistics and Linguistic Theory | VOL. 0
Doris Eveline SchönefeldDoris Eveline Schönefeld
16 Jul 2024
Corpus Linguistics and Linguistic Theory | VOL. 0

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Corpus of Regional American Language from YouTube

Abstract

Talk to us

Similar Papers

More From: Digital Humanities in the Nordic and Baltic Countries Publications