The Impact of Data Quality on Software Testing Effort Prediction

Łukasz Radliński

doi:10.3390/electronics12071656

Abstract

Background: This paper investigates the impact of data quality on the performance of models predicting effort on software testing. Data quality was reflected by training data filtering strategies (data variants) covering combinations of Data Quality Rating, UFP Rating, and a threshold of valid cases. Methods: The experiment used the ISBSG dataset and 16 machine learning models. A process of three-fold cross-validation repeated 20 times was used to train and evaluate each model with each data variant. Model performance was assessed using absolute errors of prediction. A ‘win–tie–loss’ procedure, based on the Wilcoxon signed-rank test, was applied to identify the best models and data variants. Results: Most models, especially the most accurate, performed the best on a complete dataset, even though it contained cases with low data ratings. The detailed results include the rankings of the following: (1) models for particular data variants, (2) data variants for particular models, and (3) the best-performing combinations of models and data variants. Conclusions: Arbitrary and restrictive data selection to only projects with Data Quality Rating and UFP Rating of ‘A’ or ‘B’, commonly used in the literature, does not seem justified. It is recommended not to exclude cases with low data ratings to achieve better accuracy of most predictive models for testing effort prediction.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Electronics	Publication Date: Mar 31, 2023
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

The Impact of Data Quality on Software Testing Effort Prediction

Abstract

Talk to us

Similar Papers

More From: Electronics

Lead the way for us

Similar Papers

Pushing the limits of solubility prediction via quality-oriented data selection.
Murat Cihan Sorkun ... Süleyman Er
iScience | VOL. 24
Murat Cihan Sorkun, et. al.Murat Cihan Sorkun ... Süleyman Er
17 Dec 2020
iScience | VOL. 24

The Impact of Data Quality on Neural Network Models
Chunmei Li ... Zhao Li
-
Chunmei Li, et. al.Chunmei Li ... Zhao Li
25 Apr 2019
25 Apr 2019

Data Quality and Study Compliance Among College Students Across 2 Recruitment Sources: Two Study Investigation.
Abby L Braitman ... Rachel I Macintyre
JMIR Formative Research | VOL. 6
Abby L Braitman, et. al.Abby L Braitman ... Rachel I Macintyre
09 Dec 2022
JMIR Formative Research | VOL. 6

The impact of data quality filtering of opportunistic citizen science data on species distribution model performance
Camille Van Eupen ... Stijn Luca
Ecological Modelling | VOL. 444
Camille Van Eupen, et. al.Camille Van Eupen ... Stijn Luca
02 Feb 2021
Ecological Modelling | VOL. 444

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

The Impact of Data Quality on Software Testing Effort Prediction

Abstract

Talk to us

Similar Papers

More From: Electronics