Abstract

The Seoul Corpus is a spontaneous speech corpus in Seoul Korean fully segmented with several levels of annotations in the Praat Textgrid format. A total of 40 people who were balanced for age and sex participated in the recordings. Each had an interview about various topics for an hour, and the recordings were labeled first by forced alignment using the HTK and then were fine-tuned by human labelers. About 220,000 phrasal words were included and 1,135,263 phoneme tokens were labeled. The corpus has already been distributed to the research community free of charge.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call