Multi-word Extraction Research Articles

As a hybrid of N-gram in natural language processing and collocation in statistical linguistics, multi-word is becoming a hot topic in area of text mining and information retrieval. In this paper, a study concerning distribution of multi-words is carried out to explore a theoretical basis for probabilistic term-weighting scheme. Specifically, the Poisson distribution, zero-inflated binomial distribution, and G-distribution are comparatively studied on a task of predicting probabilities of multi-words' occurrences using these distributions, for both technical multi-words and nontechnical multi-words. In addition, a rule-based multi-word extraction algorithm is proposed to extract multi-words from texts based on words' occurring patterns and syntactical structures. Our experimental results demonstrate that G-distribution has the best capability to predict probabilities of frequency of multi-words' occurrence and the Poisson distribution is comparable to zero-inflated binomial distribution in estimation of multi-word distribution. The outcome of this study validates that burstiness is a universal phenomenon in linguistic count data, which is applicable not only for individual content words but also for multi-words.

Read full abstract

One of the main themes which support text mining is text representation; that is, its task is to look for appropriate terms to transfer documents into numerical vectors. Recently, many efforts have been invested on this topic to enrich text representation using vector space model (VSM) to improve the performances of text mining techniques such as text classification and text clustering. The main concern in this paper is to investigate the effectiveness of using multi-words for text representation on the performances of text classification. Firstly, a practical method is proposed to implement the multi-word extraction from documents based on the syntactical structure. Secondly, two strategies as general concept representation and subtopic representation are presented to represent the documents using the extracted multi-words. In particular, the dynamic k-mismatch is proposed to determine the presence of a long multi-word which is a subtopic of the content of a document. Finally, we carried out a series of experiments on classifying the Reuters-21578 documents using the representations with multi-words. We used the performance of representation in individual words as the baseline, which has the largest dimension of feature set for representation without linguistic preprocessing. Moreover, linear kernel and non-linear polynomial kernel in support vector machines (SVM) are examined comparatively for classification to investigate the effect of kernel type on their performances. Index terms with low information gain (IG) are removed from the feature set at different percentages to observe the robustness of each classification method. Our experiments demonstrate that in multi-word representation, subtopic representation outperforms the general concept representation and the linear kernel outperforms the non-linear kernel of SVM in classifying the Reuters data. The effect of applying different representation strategies is greater than the effect of applying the different SVM kernels on classification performance. Furthermore, the representation using individual words outperforms any representation using multi-words. This is consistent with the major opinions concerning the role of linguistic preprocessing on documents’ features when using SVM for text classification.

Read full abstract

Multi-word Extraction Research Articles

Articles published on Multi-word Extraction

A Hybrid Multi-Word Terms Extraction System Applied to Topic Detection

Statistics and linguistic rules in multiword extraction: a comparative analysis

Clustering of PubMed abstracts using nearer terms of the domain.

DISTRIBUTION OF MULTI-WORDS IN CHINESE AND ENGLISH DOCUMENTS

Improving effectiveness of mutual information for substantival multiword expression extraction

Text classification based on multi-word with support vector machine

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Multi-word Extraction Research Articles

Articles published on Multi-word Extraction

A Hybrid Multi-Word Terms Extraction System Applied to Topic Detection

Statistics and linguistic rules in multiword extraction: a comparative analysis

Clustering of PubMed abstracts using nearer terms of the domain.

DISTRIBUTION OF MULTI-WORDS IN CHINESE AND ENGLISH DOCUMENTS

Improving effectiveness of mutual information for substantival multiword expression extraction

Text classification based on multi-word with support vector machine