Abstract

This study proposes a novel method for the segmentation of Archaic Chinese sentences based on a bidirectional long short-term memory (LSTM) + conditional random field (CRF) model. The method added a layer of linear statistical model to the traditional bidirectional LSTM neural network; it can be used for sequence annotation from the sentence level. In addition, this model introduced the stochastic gradient descent (SGD) to prevent excessive fitting, and the viterbi algorithm was used to calculate the optimal sequence of the sentences. In the experiment, this study tests the performance of the proposed method using the History of the Han Dynasty, the History of the later Han Dynasty, Three Kingdoms, and the Book of Jin, amongst others. The results show that the precision value, recall value, and F1 value are 0.77, 0.75, and 0.76, respectively, in the open test, and 0.90, 0.88, and 0.76, respectively, in the closed test.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call