A boundary-based tokenization technique for extractive text summarization

Nnaemeka M Oparauwah Nnaemeka M Oparauwah,Ikechukwu I Ayogu Ikechukwu I Ayogu,Juliet N Odii Juliet N Odii,Vitalis C Iwuchukwu Vitalis C Iwuchukwu

doi:10.30574/wjarr.2021.11.2.0351

Nnaemeka M Oparauwah Nnaemeka M Oparauwah, Ikechukwu I Ayogu Ikechukwu I Ayogu + Show 2 more

Open Access

https://doi.org/10.30574/wjarr.2021.11.2.0351

Copy DOI

Journal: World Journal of Advanced Research and Reviews	Publication Date: Aug 30, 2021
Citations: 1	License type: cc-by-nc-sa

Abstract

The need to extract and manage vital information contained in copious volumes of text documents has given birth to several automatic text summarization (ATS) approaches. ATS has found application in academic research, medical health records analysis, content creation and search engine optimization, finance and media. This study presents a boundary-based tokenization method for extractive text summarization. The proposed method performs word tokenization by defining word boundaries in place of specific delimiters. An extractive summarization algorithm was further developed based on the proposed boundary-based tokenization method, as well as word length consideration to control redundancy in summary output. Experimental results showed that the proposed approach enhanced word tokenization by enhancing the selection of appropriate keywords from text document to be used for summarization.

Full Text