Machine learning based software effort estimation using development-centric features for crowdsourcing platform

Anum Yasmin,Ameen Banjar,Ali Daud,Wasi Haider

doi:10.3233/ida-237366

Abstract

Crowd-Sourced software development (CSSD) is getting a good deal of attention from the software and research community in recent times. One of the key challenges faced by CSSD platforms is the task selection mechanism which in practice, contains no intelligent scheme. Rather, rule-of-thumb or intuition strategies are employed, leading to biasness and subjectivity. Effort considerations on crowdsourced tasks can offer good foundation for task selection criteria but are not much investigated. Software development effort estimation (SDEE) is quite prevalent domain in software engineering but only investigated for in-house development. For open-sourced or crowdsourced platforms, it is rarely explored. Moreover, Machine learning (ML) techniques are overpowering SDEE with a claim to provide more accurate estimation results. This work aims to conjoin ML-based SDEE to analyze development effort measures on CSSD platform. The purpose is to discover development-oriented features for crowdsourced tasks and analyze performance of ML techniques to find best estimation model on CSSD dataset. TopCoder is selected as target CSSD platform for the study. TopCoder’s development tasks data with development-centric features are extracted, leading to statistical, regression and correlation analysis to justify features’ significance. For effort estimation, 10 ML families with 2 respective techniques are applied to get broader aspect of estimation. Five performance metrices (MSE, RMSE, MMRE, MdMRE, Pred (25) and Welch’s statistical test are incorporated to judge the worth of effort estimation model’s performance. Data analysis results show that selected features of TopCoder pertain reasonable model significance, regression, and correlation measures. Findings of ML effort estimation depicted that best results for TopCoder dataset can be acquired by linear, non-linear regression and SVM family models. To conclude, the study identified the most relevant development features for CSSD platform, confirmed by in-depth data analysis. This reflects careful selection of effort estimation features to offer good basis of accurate ML estimate.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Machine learning based software effort estimation using development-centric features for crowdsourcing platform

Abstract

Talk to us

Similar Papers

More From: Intelligent Data Analysis

Lead the way for us

Similar Papers

Study of Learning Techniques for Effort Estimation in Object-Oriented Software Development
Suyash Shukla ... Sandeep Kumar
IEEE Transactions on Engineering Management | VOL. 71
Suyash Shukla, et. al.Suyash Shukla ... Sandeep Kumar
01 Jan 2024
IEEE Transactions on Engineering Management | VOL. 71

Predictive analytics approaches for software effort estimation: A review
A G Priya Varshini
Indian Journal of Science and Technology | VOL. 13
A G Priya VarshiniA G Priya Varshini
05 Jun 2020
Indian Journal of Science and Technology | VOL. 13

The Implementation of Machine Learning for Software Effort Estimation: A Literature Review
Eva Hariyanti ... Fakhrana Almas Syah Yahrani
Khazanah Informatika : Jurnal Ilmu Komputer dan Informatika | VOL. 10
Eva Hariyanti, et. al.Eva Hariyanti ... Fakhrana Almas Syah Yahrani
30 Apr 2024
Khazanah Informatika : Jurnal Ilmu Komputer dan Informatika | VOL. 10

Evaluating the Progressive Performance of Machine Learning Techniques on E-commerce Data
Bindu Madhuri Cheekati ... Sai Varun Padala
-
Bindu Madhuri Cheekati, et. al.Bindu Madhuri Cheekati ... Sai Varun Padala
29 Oct 2017
29 Oct 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Machine learning based software effort estimation using development-centric features for crowdsourcing platform

Abstract

Talk to us

Similar Papers

More From: Intelligent Data Analysis