Bug Prediction Using Source Code Embedding Based on Doc2Vec

Tamás Aladics,Judit Jász,Rudolf Ferenc

doi:10.1007/978-3-030-87007-2_27

Abstract

Bug prediction is a resource demanding task that is hard to automate using static source code analysis. In many fields of computer science, machine learning has proven to be extremely useful in tasks like this, however, for it to work we need a way to use source code as input. We propose a simple, but meaningful representation for source code based on its abstract syntax tree and the Doc2Vec embedding algorithm. This representation maps the source code to a fixed length vector which can be used for various upstream tasks – one of which is bug prediction. We measured this approach’s validity by itself and its effectiveness compared to bug prediction based solely on code metrics. We also experimented on numerous machine learning approaches to check the connection between different embedding parameters with different machine learning models. Our results show that this representation provides meaningful information as it improves the bug prediction accuracy in most cases, and is always at least as good as only using code metrics as features.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Bug Prediction Using Source Code Embedding Based on Doc2Vec

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Enhanced Bug Prediction in JavaScript Programs with Hybrid Call-Graph Based Invocation Metrics
Gábor Antal ... Péter Hegedűs
Technologies | VOL. 9
Gábor Antal, et. al.Gábor Antal ... Péter Hegedűs
30 Dec 2020
Technologies | VOL. 9

Method level static source code analysis on behavioral change impact analysis in software regression testing
Fredrick Mugambi Muthengi ... Faith Mueni Musyoka
Indonesian Journal of Electrical Engineering and Computer Science | VOL. 35
Fredrick Mugambi Muthengi, et. al.Fredrick Mugambi Muthengi ... Faith Mueni Musyoka
01 Jul 2024
Indonesian Journal of Electrical Engineering and Computer Science | VOL. 35

Rule based graph visualization for software systems
Tibor Brunner ... Máté Cserép
-
Tibor Brunner, et. al.Tibor Brunner ... Máté Cserép
01 Jan 2015
01 Jan 2015

КОМБІНОВАНІ ПІДХОДИ ДО СТАТИЧНОГО АНАЛІЗУ КОДУ З ВИКОРИСТАННЯМ НЕЙРОННИХ МЕРЕЖ
Illia Vokhranov ... Bogdan Bulakh
Інфокомунікаційні та комп’ютерні технології | VOL. -
Illia Vokhranov, et. al.Illia Vokhranov ... Bogdan Bulakh
01 Jan 2023
Інфокомунікаційні та комп’ютерні технології | VOL. -

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Bug Prediction Using Source Code Embedding Based on Doc2Vec

Abstract

Talk to us

Similar Papers