Imperfect Transcripts Research Articles

This paper describes our work in semantic interpretation of a “multimodal language” with speech and gestures using latent semantic analysis (LSA). Our aim is to infer the domain-specific informational goal of multimodal inputs. The informational goal is characterized by lexical terms used in the spoken modality, partial semantics of gestures in the pen modality, as well as term co-occurrence patterns across modalities, leading to “multimodal terms.” We designed and collected a multimodal corpus of navigational inquiries. We also obtained perfect (i.e. manual) and imperfect (i.e. automatic via recognition) transcriptions for these. We automatically align parsed spoken locative references (SLRs) with their corresponding pen gesture(s) using the Viterbi alignment, according to their numeric and location type features. Then, we characterize each cross-modal integration pattern as a 3-tuple multimodal term with SLR, pen gesture type and their temporal relationship. We propose to use latent semantic analysis (LSA) to derive the latent semantics from manual (i.e. perfect) and automatic (i.e. imperfect) transcriptions of the collected multimodal inputs. In order to achieve this, both multimodal and lexical terms are used to compose an inquiry-term matrix, which is then factorized using singular value decomposition (SVD) to derive the latent semantics automatically. Informational goal inference based on the latent semantics shows that the informational goal inference accuracy of a disjoint test set is 99% and 84% when a perfect and imperfect projection model is used respectively, which performs significantly better than (at least 9.9% absolute) the baseline performance using vector-space model (VSM).

Read full abstract

In this paper, we propose a one step rhetorical structure parsing, chunking and extractive summarization approach to automatically generate meeting minutes from parliamentary speech using acoustic and lexical features. We investigate how to use lexical features extracted from imperfect ASR transcriptions, together with acoustic features extracted from the speech itself, to form extractive summaries with the structure of meeting minutes. Each business item in the minute is modeled as a rhetorical chunk which consists of smaller rhetorical units. Principal Component Analysis (PCA) graphs of both acoustic and lexical features in meeting speech show clear self-clustering of speech utterances according to the underlying rhetorical state-for example acoustic and lexical feature vectors from the question and answer or motion of a parliamentary speech, are grouped together. We then propose a Conditional Random Fields (CRF)-based approach to perform both rhetorical structure modeling and extractive summarization in one step, by chunking, parsing and extraction of salient utterances. Extracted salient utterances are grouped under the labels of each rhetorical state, emulating meeting minutes to yield summaries that are more easily understandable by humans. We compare this approach to different machine learning methods. We show that our proposed CRF-based one step minute generation system obtains the best summarization performance both in terms of ROUGE-L F-measure at 74.5% and by human evaluation, at 77.5% on average.

Read full abstract

Imperfect Transcripts Research Articles

Related Topics

Articles published on Imperfect Transcripts

ALISA: An automatic lightly supervised speech segmentation and alignment tool

Latent Semantic Analysis for Multimodal User Input With Speech and Gestures

Automatic Parliamentary Meeting Minute Generation Using Rhetorical Structure Modeling

Integrating imperfect transcripts into speech recognition systems for building high-quality corpora

Question answering from lecture videos based on an automatic semantic annotation

Searching large collections of recorded speech: A preliminary study

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Imperfect Transcripts Research Articles

Related Topics

Articles published on Imperfect Transcripts

ALISA: An automatic lightly supervised speech segmentation and alignment tool

Latent Semantic Analysis for Multimodal User Input With Speech and Gestures

Automatic Parliamentary Meeting Minute Generation Using Rhetorical Structure Modeling

Integrating imperfect transcripts into speech recognition systems for building high-quality corpora

Question answering from lecture videos based on an automatic semantic annotation

Searching large collections of recorded speech: A preliminary study