Measuring Dynamic Correlations of Words in Written Texts with an Autocorrelation Function

Hiroshi Ogura,Masato Kondo,Hiromi Amano

doi:10.4236/jdaip.2019.72004

Hiroshi Ogura, Masato Kondo + Show 1 more

Open Access

https://doi.org/10.4236/jdaip.2019.72004

Copy DOI

Abstract

In this study, we regard written texts as time series data and try to investigate dynamic correlations of word occurrences by utilizing an autocorrelation function (ACF). After defining appropriate formula for the ACF that is suitable for expressing the dynamic correlations of words, we use the formula to calculate ACFs for frequent words in 12 books. The ACFs obtained can be classified into two groups: One group of ACFs shows dynamic correlations, with these ACFs well described by a modified Kohlrausch-Williams-Watts (KWW) function; the other group of ACFs shows no correlations, with these ACFs fitted by a simple stepdown function. A word having the former ACF is called a Type-I word and a word with the latter ACF is called a Type-II word. It is also shown that the ACFs of Type-II words can be derived theoretically by assuming that the stochastic process governing word occurrence is a homogeneous Poisson point process. Based on the fitting of the ACFs by KWW and stepdown functions, we propose a measure of word importance which expresses the extent to which a word is important in a particular text. The validity of the measure is confirmed by using the Kleinburg’s burst detection algorithm.

Highlights

We use language to convey our ideas
The autocorrelation function (ACF) obtained can be classified into two groups: One group of ACFs shows dynamic correlations, with these ACFs well described by a modified Kohlrausch-Williams-Watts (KWW) function; the other group of ACFs shows no correlations, with these ACFs fitted by a simple stepdown function
Starting from the standard definition of an ACF in the signal processing area, we derived a normalized expression for an ACF that is suitable to express the dynamic correlation of word occurrences

Summary

Introduction

We use language to convey our ideas. Since our physical function is limited to speaking or writing only one word at a time, we must transform our complex ideas into linear strings of words. This approach has been successfully applied to the extraction of semantic representations [1], automatic key word and key phrase extraction [2] [3], local or global context analysis [4], measuring similarities at the word or context level [5], and many other tasks Another way to investigate correlations in linguistic data is to use a mapping scheme, that is, to translate the given sequence of words or characters in a text into a time series and thereby capture the correlations in a dynamical framework. The goal of this study is to find a modification of the word-level mapping that is suitable for defining and calculating appropriate ACFs in the mapping scheme

Ogura et al DOI

Models of Word Occurrences

Models of Linguistic Data with ACFs

Calculation of ACF for Written Texts

Typical Examples of Correlated and Non-Correlated ACFs

Curve Fitting Using Model Functions

Classification of Frequent Words

Model Selection Using the Bayesian Information Criterion

Stochastic Model for Type-II Words

Measure of Dynamic Correlation

Conclusions

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Data Analysis and Information Processing	Publication Date: Jan 1, 2019
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Measuring Dynamic Correlations of Words in Written Texts with an Autocorrelation Function

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Data Analysis and Information Processing

Lead the way for us

Similar Papers

Origin of Dynamic Correlations of Words in Written Texts
Hiroshi Ogura ... Masato Kondo
Journal of Data Analysis and Information Processing | VOL. 07
Hiroshi Ogura, et. al.Hiroshi Ogura ... Masato Kondo
01 Jan 2019
Journal of Data Analysis and Information Processing | VOL. 07

Simulation of pseudo-text synthesis for generating words with long-range dynamic correlations
Hiroshi Ogura ... Hiromi Amano
SN Applied Sciences | VOL. 2
Hiroshi Ogura, et. al.Hiroshi Ogura ... Hiromi Amano
16 Jul 2020
SN Applied Sciences | VOL. 2

Significance of coincident spiking considering inter-spike interval variability and serial interval correlation
Grün Sonja
Frontiers in Computational Neuroscience | VOL. 2
Grün SonjaGrün Sonja
01 Jan 2008
Frontiers in Computational Neuroscience | VOL. 2

A stochastic model of word occurrences in hierarchically structured written texts
Hiroshi Ogura ... Hiromi Amano
SN Applied Sciences | VOL. 4
Hiroshi Ogura, et. al.Hiroshi Ogura ... Hiromi Amano
14 Feb 2022
SN Applied Sciences | VOL. 4

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Measuring Dynamic Correlations of Words in Written Texts with an Autocorrelation Function

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Journal of Data Analysis and Information Processing