Joint dialog act segmentation and recognition in human conversations using attention to dialog context

Tianyu Zhao,Tatsuya Kawahara

doi:10.1016/j.csl.2019.03.001

Tianyu Zhao, Tatsuya Kawahara

Open Access

PDF Available

https://doi.org/10.1016/j.csl.2019.03.001

Copy DOI

Export

Save

Cite

Journal: Computer Speech & Language	Publication Date: Mar 15, 2019
Citations: 8	License type: publisher-specific-oa

Affiliation: Kyoto University

Abstract
Full-Text PDF
Similar Papers

Abstract

Listen

A dialog act represents the communicative function of an utterance in a conversation, and thus provides informative cues for understanding, managing, and generating dialog. While most spoken dialog systems process user input and system output at the turn level, a single turn can consist of multiple dialog acts in human conversations. Therefore, segmenting turn-level tokens into a meaningful dialog act unit is just as important as recognizing the dialog act. Towards joint segmentation and recognition of dialog acts, we propose an encoder–decoder model featuring joint coding and incorporate contextual information by means of an attentional mechanism. The proposed encoder–decoder outperforms other models in segmentation, and the application of attentions significantly reduces recognition error rates. By combining the encoder–decoder model with contextual attention, we achieve state-of-the-art performance in the joint evaluation of dialog act segmentation and recognition.

Full Text