Scaling out for extreme scale corpus data

Matthew Coole,Paul Rayson,John Mariani

doi:10.1109/bigdata.2015.7363933

Abstract

Much of the previous work in Big Data has focussed on numerical sources of information. However, with the ‘narrative turn’ in many disciplines gathering pace and commercial organisations beginning to realise the value of their textual assets, natural language data is fast catching up as an exploitable source of information for decision making. With vast quantities of unstructured textual data on the web, in social media, and in newly digitised historical document archives, the 5Vs (Volume, Velocity, Variety, Value and Veracity) apply equally well, if not more so, to big textual data. Corpus linguistics, the computer-aided study of large collections of naturally occurring language data, has been dealing with big data for fifty years. Corpus linguistics methods impose complex requirements on the retrieval, annotation and analysis of text in terms of displaying narrow contexts for each occurrence of a word or linguistic feature being studied and counting co-occurrences with other words or features to determine significant patterns in language. This, coupled with the distribution of language features in accordance with Zipf's Law, poses complex challenges for data models and corpus software dealing with extreme scale language data. A related issue is the non-random nature of language and the ‘burstiness’ of word occurrences, or what we might put in Big Data terms as a sixth ‘V’ called Viscosity. We report experiments to examine and compare the capabilities of two No-SQL databases in clustered configurations for the indexing, retrieval and analysis of billion-word corpora, since this size is the current state-of-the-art in corpus linguistics. We find that modern DBMSs (Database Management Systems) are capable of handling this extreme scale corpus data set for simple queries but are limited when querying for more frequent words or more complex queries.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Scaling out for extreme scale corpus data

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

A time-sensitive historical thesaurus-based semantic tagger for deep semantic annotation
Scott Piao ... Marc Alexander
Computer Speech & Language | VOL. 46
Scott Piao, et. al.Scott Piao ... Marc Alexander
17 May 2017
Computer Speech & Language | VOL. 46

Big Data Storage and Management: Challenges and Opportunities
Jaroslav Pokorný
-
Jaroslav PokornýJaroslav Pokorný
01 Jan 2017
01 Jan 2017

An adaptive spark-based framework for querying large-scale NoSQL and relational databases.
Eman Khashan ... Ali Eldesouky
PloS one | VOL. 16
Eman Khashan, et. al.Eman Khashan ... Ali Eldesouky
19 Aug 2021
PloS one | VOL. 16

Semantic Tagging for the Urdu Language: Annotated Corpus and Multi-Target Classification Methods
Jawad Shafi ... Paul Rayson
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 22
Jawad Shafi, et. al.Jawad Shafi ... Paul Rayson
17 Jun 2023
ACM Transactions on Asian and Low-Resource Language Information Processing | VOL. 22

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Scaling out for extreme scale corpus data

Abstract

Talk to us

Similar Papers