New Word Extraction From Chinese Financial Documents

Liwei Yan,Bo Bai,Wei Chen,Dapeng Oliver Wu

doi:10.1109/lsp.2017.2690599

New Word Extraction From Chinese Financial Documents

Liwei Yan, Bo Bai + Show 2 more

https://doi.org/10.1109/lsp.2017.2690599

Copy DOI

Journal: IEEE Signal Processing Letters	Publication Date: Jun 1, 2017
Citations: 7

Affiliation: Tsinghua University, Huawei Technologies (China), University of Florida

#Word Extraction #Natural Language Processing + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

With the tremendous development of data science, using unstructured documents to analyze marketing dynamics is attracting a great deal of attention. In this letter, we propose an iterative scheme to extract the new words, which is often a bottleneck for Chinese natural language processing (NLP) in financial markets analysis. In contrast to existing static features, the key novelty is the proposed dynamic features that characterize the similarity of context patterns. Via iteration, distinguishable seed context patterns are extracted. Tested on a 203 MB corpus, 19 291 words representing emerging industries, entities, projects, and products were extracted with a precision of 89.8% and recall of 88.9%, which outperforms most competitor methods.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: IEEE Signal Processing Letters

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.