Authorship Attribution via Network Motifs Identification

Vanessa Queiroz Marinho,Diego Raphael Amancio,Graeme Hirst

doi:10.1109/bracis.2016.071

Abstract

Concepts and methods of complex networks can be used to analyse texts at their different complexity levels. Examples of natural language processing (NLP) tasks studied via topological analysis of networks are keyword identification, automatic extractive summarization and authorship attribution. Even though a myriad of network measurements have been applied to study the authorship attribution problem, the use of motifs for text analysis has been restricted to a few works. The goal of this paper is to apply the concept of motifs, recurrent interconnection patterns, in the authorship attribution task. The absolute frequencies of all thirteen directed motifs with three nodes were extracted from the co-occurrence networks and used as classification features. The effectiveness of these features was verified with four machine learning methods. The results show that motifs are able to distinguish the writing style of different authors. In our best scenario, 57.5% of the books were correctly classified. The chance baseline for this problem is 12.5%. In addition, we have found that function words play an important role in these recurrent patterns. Taken together, our findings suggest that motifs should be further explored in other related linguistic tasks.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Authorship Attribution via Network Motifs Identification

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Corpus of transcribed parliamentary speeches for authorship attribution and author profiling tasks
Jurgita Kapočiūtė-Dzikienė ... Ligita Šarkutė
Kalbotyra | VOL. 66
Jurgita Kapočiūtė-Dzikienė, et. al.Jurgita Kapočiūtė-Dzikienė ... Ligita Šarkutė
30 Mar 2016
Kalbotyra | VOL. 66

Automatic authorship attribution in Albanian texts.
Arta Misini ... Endrit Fetahi
PloS one | VOL. 19
Arta Misini, et. al.Arta Misini ... Endrit Fetahi
22 Oct 2024
PloS one | VOL. 19

Relevance of Named Entities in Authorship Attribution
Germán Ríos-Toledo ... Liliana Chanona-Hernández
-
Germán Ríos-Toledo, et. al.Germán Ríos-Toledo ... Liliana Chanona-Hernández
01 Jan 2017
01 Jan 2017

Phraseology and Style in Subgenres of the Novel: A Synthesis of Corpus and Literary Perspectives
Rundong Wang ... Hongwei Zhan
Style | VOL. 56
Rundong Wang, et. al.Rundong Wang ... Hongwei Zhan
01 Aug 2022
Style | VOL. 56

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Authorship Attribution via Network Motifs Identification

Abstract

Talk to us

Similar Papers