Quality assurance strategies for machine learning applications in big data analytics: an overview

Mihajlo Ogrizović,Dražen Drašković,Dragan Bojić

doi:10.1186/s40537-024-01028-y

Abstract

Machine learning (ML) models have gained significant attention in a variety of applications, from computer vision to natural language processing, and are almost always based on big data. There are a growing number of applications and products with built-in machine learning models, and this is the area where software engineering, artificial intelligence and data science meet. The requirement for a system to operate in a real-world environment poses many challenges, such as how to design for wrong predictions the model may make; How to assure safety and security despite possible mistakes; which qualities matter beyond a model’s prediction accuracy; How can we identify and measure important quality requirements, including learning and inference latency, scalability, explainability, fairness, privacy, robustness, and safety. It has become crucial to test thoroughly these models to assess their capabilities and potential errors. Existing software testing methods have been adapted and refined to discover faults in machine learning and deep learning models. This paper covers a taxonomy, a methodologically uniform presentation of all presented solutions to the aforementioned issues, as well as conclusions about possible future development trends. The main contributions of this paper are a classification that closely follows the structure of the ML-pipeline, a precisely defined role of each team member within that pipeline, an overview of trends and challenges in the combination of ML and big data analytics, with uses in the domains of industry and education.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Quality assurance strategies for machine learning applications in big data analytics: an overview

Abstract

Talk to us

Similar Papers

More From: Journal of Big Data

Lead the way for us

Journal: Journal of Big Data	Publication Date: Oct 30, 2024
License type: CC BY-NC-ND 4.0

Similar Papers

A novel deep-learning technique for forecasting oil price volatility using historical prices of five precious metals in context of green financing – A comparison of deep learning, machine learning, and statistical models
Muhammad Mohsin ... Fouad Jamaani
Resources Policy | VOL. 86
Muhammad Mohsin, et. al.Muhammad Mohsin ... Fouad Jamaani
01 Oct 2023
Resources Policy | VOL. 86

Deep learning‐based smishing message identification using regular expression feature generation
Aakanksha Sharaff ... Vrihas Pathak
Expert Systems | VOL. 40
Aakanksha Sharaff, et. al.Aakanksha Sharaff ... Vrihas Pathak
05 Oct 2022
Expert Systems | VOL. 40

Explainable artificial intelligence (XAI) for predicting the need for intubation in methanol-poisoned patients: a study comparing deep and machine learning models
Khadijeh Moulaei ... Peyman Erfan Talab Evini
Scientific Reports | VOL. 14
Khadijeh Moulaei, et. al.Khadijeh Moulaei ... Peyman Erfan Talab Evini
08 Jul 2024
Scientific Reports | VOL. 14

EESD special issue: AI and data‐driven methods in earthquake engineering – (Part 1)
Xinzheng Lu ... Henry Burton
Earthquake Engineering & Structural Dynamics | VOL. 52
Xinzheng Lu, et. al.Xinzheng Lu ... Henry Burton
04 May 2023
Earthquake Engineering & Structural Dynamics | VOL. 52

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Quality assurance strategies for machine learning applications in big data analytics: an overview

Abstract

Talk to us

Similar Papers

More From: Journal of Big Data