A Comparative Study of Data Collection Periods for Just-In-Time Defect Prediction Using the Automatic Machine Learning Method

Kosuke Ohara,Minoru Kawahara,Hirohisa Aman,Tomoyuki Yokogawa,Sousuke Amasaki

doi:10.1587/transinf.2022mpl0002

A Comparative Study of Data Collection Periods for Just-In-Time Defect Prediction Using the Automatic Machine Learning Method

Kosuke Ohara, Minoru Kawahara + Show 3 more

Open Access

https://doi.org/10.1587/transinf.2022mpl0002

Copy DOI

Journal: IEICE Transactions on Information and Systems	Publication Date: Feb 1, 2023
License type: free

Affiliation: Ehime University, Center for Information Technology, Okayama Prefectural University

#Defect Prediction Models #Selection Of Machine Learning Algorithms + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

This paper focuses on the “data collection period” for training a better Just-In-Time (JIT) defect prediction model — the early commit data vs. the recent one —, and conducts a large-scale comparative study to explore an appropriate data collection period. Since there are many possible machine learning algorithms for training defect prediction models, the selection of machine learning algorithms can become a threat to validity. Hence, this study adopts the automatic machine learning method to mitigate the selection bias in the comparative study. The empirical results using 122 open-source software projects prove the trend that the dataset composed of the recent commits would become a better training set for JIT defect prediction models.

Full Text