Chinese POS Tagging Research Articles

Chinese natural language processing tasks often require the solution of Chinese word segmentation and POS tagging problems. Traditional Chinese word segmentation and POS tagging methods mainly use simple matching algorithms based on lexicons and rules. The simple matching or statistical analysis requires manual word segmentation followed by POS tagging, which leads to the inability to meet the practical requirements for label prediction accuracy. With the continuous development of deep learning technology, data-driven machine learning models provide new opportunities for automated Chinese word segmentation and POS tagging. Therefore, a data-driven automated Chinese word segmentation and POS tagging model is proposed in order to address the above problems. Firstly, the main idea and overall framework of the proposed automated model are outlined, and the tagging strategy and neural network language model used are described. Secondly, two main optimisations are made on the input side of the model: (1) the use of word2Vec for the representation of text features, thus representing the text as a distributed word vector; and (2) the use of an improved AlexNet for efficient encoding of long-range word, and the addition of an attention mechanism to the model. Finally, on the output side, an additional auxiliary loss function was designed to optimise the Chinese text based on its frequency. The experimental results show that the proposed model can significantly improve the accuracy and operational efficiency of Chinese word segmentation and POS tagging compared with other existing models, thus verifying its effectiveness and advancement.

Read full abstract

Words provide a useful source of information for Chinese NLP, and word segmentation has been taken as a pre-processing step for most downstream tasks. For many NLP tasks, however, word segmentation can introduce noise and lead to error propagation. The rise of neural representation learning models allows sentence-level semantic information to be collected from characters directly. As a result, it is an empirical question whether a fully character-based model should be used instead of first performing word segmentation. We investigate a neural representation that simultaneously encodes character and word information without the need for segmentation. In particular, candidate words are found in a sentence by matching with a pre-defined lexicon. A lattice structured LSTM is used to encode the resulting word-character lattice, where gate vectors are used to control information flow through words, so that the more useful words can be automatically identified by end-to-end training. We compare the performance of the resulting lattice LSTM and baseline sequence LSTM structures over both character sequences and automatically segmented word sequences. Results on NER show that the character-word lattice model can significantly improve the performance. In addition, as a general sentence representation architecture, character-word lattice LSTM can also be used for learning contextualized representations. To this end, we compare lattice LSTM structure with its sequential LSTM counterpart, namely ELMo. Results show that our lattice version of ELMo gives better language modeling performances. On Chinese POS-tagging, chunking and syntactic parsing tasks, the resulting contextualized Chinese embeddings also give better performance than ELMo trained on the same data.

Read full abstract

Chinese POS Tagging Research Articles

Related Topics

Articles published on Chinese POS Tagging

A Data-Driven Model for Automated Chinese Word Segmentation and POS Tagging.

Encoding multi-granularity structural information for joint Chinese word segmentation and POS tagging

Lattice LSTM for Chinese Sentence Representation

A Simple and Effective Neural Model for Joint Word Segmentation and POS Tagging

Towards Accurate and Efficient Chinese Part-of-Speech Tagging

Simple Semi-supervised Learning for Chinese Word Segmentation and Pos Tagging

Joint Chinese Word Segmentation and POS Tagging System with Undirected Graphical Models

Joint Chinese Word Segmentation and POS Tagging Using an Error-Driven Word-Character Hybrid Model

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Chinese POS Tagging Research Articles

Related Topics

Articles published on Chinese POS Tagging

A Data-Driven Model for Automated Chinese Word Segmentation and POS Tagging.

Encoding multi-granularity structural information for joint Chinese word segmentation and POS tagging

Lattice LSTM for Chinese Sentence Representation

A Simple and Effective Neural Model for Joint Word Segmentation and POS Tagging

Towards Accurate and Efficient Chinese Part-of-Speech Tagging

Simple Semi-supervised Learning for Chinese Word Segmentation and Pos Tagging

Joint Chinese Word Segmentation and POS Tagging System with Undirected Graphical Models

Joint Chinese Word Segmentation and POS Tagging Using an Error-Driven Word-Character Hybrid Model