Enhancing Software Effort Estimation with Pre-Trained Word Embeddings: A Small-Dataset Solution for Accurate Story Point Prediction

Issa Atoum,Ahmed Ali Otoom

doi:10.3390/electronics13234843

Abstract

Traditional software effort estimation methods, such as term frequency–inverse document frequency (TF-IDF), are widely used due to their simplicity and interpretability. However, they struggle with limited datasets, fail to capture intricate semantics, and suffer from dimensionality, sparsity, and computational inefficiency. This study used pre-trained word embeddings, including FastText and GPT-2, to improve estimation accuracy in such cases. Seven pre-trained models were evaluated for their ability to effectively represent textual data, addressing the fundamental limitations of TF-IDF through contextualized embeddings. The results show that combining FastText embeddings with support vector machines (SVMs) consistently outperforms traditional approaches, reducing the mean absolute error (MAE) by 5–18% while achieving accuracy comparable to deep learning models like GPT-2. This approach demonstrated the adaptability of pre-trained embeddings for small datasets, balancing semantic richness with computational efficiency. The proposed method optimized project planning and resource allocation while enhancing software development through accurate story point prediction while safeguarding privacy and security through data anonymization. Future research will explore task-specific embeddings tailored to software engineering domains and investigate how dataset characteristics, such as cultural variations, influence model performance, ensuring the development of adaptable, robust, and secure machine learning models for diverse contexts.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Enhancing Software Effort Estimation with Pre-Trained Word Embeddings: A Small-Dataset Solution for Accurate Story Point Prediction

Abstract

Talk to us

Similar Papers

More From: Electronics

Lead the way for us

Journal: Electronics	Publication Date: Dec 8, 2024
License type: CC BY 4.0

Similar Papers

Neural Machine Translation for Kashmiri to English and Hindi using Pre-trained Embeddings
Shailashree K Sheshadri ... Deepa Gupta
-
Shailashree K Sheshadri, et. al.Shailashree K Sheshadri ... Deepa Gupta
01 Dec 2022
01 Dec 2022

Evaluating shallow and deep learning strategies for the 2018 n2c2 shared task on clinical text classification.
Michel Oleynik ... Amila Kugic
Journal of the American Medical Informatics Association | VOL. 26
Michel Oleynik, et. al.Michel Oleynik ... Amila Kugic
12 Sep 2019
Evaluating shallow and deep learning strategies for the 2018 n2c2 shared task on clinical text classification.
Michel Oleynik ... Amila Kugic

Autoencoding Improves Pre-trained Word Embeddings
Masahiro Kaneko ... Danushka Bollegala
-
Masahiro Kaneko, et. al.Masahiro Kaneko ... Danushka Bollegala
01 Jan 2020
01 Jan 2020

Autoencoding Improves Pre-trained Word Embeddings
...
-
, et. al. ...
25 Nov 2020
25 Nov 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Enhancing Software Effort Estimation with Pre-Trained Word Embeddings: A Small-Dataset Solution for Accurate Story Point Prediction

Abstract

Talk to us

Similar Papers

More From: Electronics