Improving fraud detection with semi-supervised topic modeling and keyword integration.

Marco Sánchez,Luis Urquiza

doi:10.7717/peerj-cs.1733

Abstract

Fraud detection through auditors' manual review of accounting and financial records has traditionally relied on human experience and intuition. However, replicating this task using technological tools has represented a challenge for information security researchers. Natural language processing techniques, such as topic modeling, have been explored to extract information and categorize large sets of documents. Topic modeling, such as latent Dirichlet allocation (LDA) or non-negative matrix factorization (NMF), has recently gained popularity for discovering thematic structures in text collections. However, unsupervised topic modeling may not always produce the best results for specific tasks, such as fraud detection. Therefore, in the present work, we propose to use semi-supervised topic modeling, which allows the incorporation of specific knowledge of the study domain through the use of keywords to learn latent topics related to fraud. By leveraging relevant keywords, our proposed approach aims to identify patterns related to the vertices of the fraud triangle theory, providing more consistent and interpretable results for fraud detection. The model's performance was evaluated by training with several datasets and testing it with another one that did not intervene in its training. The results showed efficient performance averages with a 7% increase in performance compared to a previous job. Overall, the study emphasizes the importance of deepening the analysis of fraud behaviors and proposing strategies to identify them proactively.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Improving fraud detection with semi-supervised topic modeling and keyword integration.

Abstract

Talk to us

Similar Papers

More From: PeerJ. Computer science

Lead the way for us

Journal: PeerJ. Computer science	Publication Date: Jan 15, 2024
License type: CC BY 4.0

Similar Papers

Urdu Documents Clustering with Unsupervised and Semi-Supervised Probabilistic Topic Modeling
Mubashar Mustafa ... Hussain Ghulam
Information | VOL. 11
Mubashar Mustafa, et. al.Mubashar Mustafa ... Hussain Ghulam
05 Nov 2020
Information | VOL. 11

Sentiment-Aspect Analysis through Semi-Supervised Topic Modeling
Yong Heng Chen ... Yaojin Lin
International Journal of Database Theory and Application | VOL. 8
Yong Heng Chen, et. al.Yong Heng Chen ... Yaojin Lin
31 Dec 2015
International Journal of Database Theory and Application | VOL. 8

Product aspect extraction supervised with online domain knowledge
Tao Wang ... Huaqing Min
Knowledge-Based Systems | VOL. 71
Tao Wang, et. al.Tao Wang ... Huaqing Min
20 Jun 2014
Knowledge-Based Systems | VOL. 71

Discovering the Research Topics on Construction Safety and Health Using Semi-Supervised Topic Modeling
Kai Zhou ... Jianli Chen
Buildings | VOL. 13
Kai Zhou, et. al.Kai Zhou ... Jianli Chen
28 Apr 2023
Buildings | VOL. 13

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Improving fraud detection with semi-supervised topic modeling and keyword integration.

Abstract

Talk to us

Similar Papers

More From: PeerJ. Computer science