Automatic data acquisition for deep learning

Jiabin Liu,Fu Zhu,Nan Tang,Yuyu Luo,Chengliang Chai

doi:10.14778/3476311.3476333

Abstract

Deep learning (DL) has widespread applications and has revolutionized many industries. Although automated machine learning (AutoML) can help us away from coding for DL models, the acquisition of lots of high-quality data for model training remains a main bottleneck for many DL projects, simply because it requires high human cost. Despite many works on weak supervision ( i.e. , adding weak labels to seen data) and data augmentation ( i.e. , generating more data based on seen data), automatically acquiring training data, via smartly searching a pool of training data collected from open ML benchmarks and data markets, is not explored. In this demonstration, we demonstrate a new system, automatic data acquisition (AutoData), which automatically searches training data from a heterogeneous data repository and interacts with AutoML. It faces two main challenges. (1) How to search high-quality data from a large repository for a given DL task? (2) How does AutoData interact with AutoML to guide the search? To address these challenges, we propose a reinforcement learning (RL)-based framework in AutoData to guide the iterative search process. AutoData encodes current training data and feedbacks of AutoML, learns a policy to search fresh data, and trains in iterations. We demonstrate with two real-life scenarios, image classification and relational data prediction, showing that AutoData can select high-quality data to improve the model.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Automatic data acquisition for deep learning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment

Lead the way for us

Journal: Proceedings of the VLDB Endowment	Publication Date: Jul 1, 2021
Citations: 12

Similar Papers

Deep Learning with Data Augmentation to Add Data Around Classification Boundaries
Hideki Fujinami ... Masayuki Goto
Industrial Engineering & Management Systems | VOL. 20
Hideki Fujinami, et. al.Hideki Fujinami ... Masayuki Goto
30 Sep 2021
Industrial Engineering & Management Systems | VOL. 20

Clinically Relevant Vulnerabilities of Deep Machine Learning Systems for Skin Cancer Diagnosis
Xinyi Du-Harpur ... Magnus D Lynch
Journal of Investigative Dermatology | VOL. 141
Xinyi Du-Harpur, et. al.Xinyi Du-Harpur ... Magnus D Lynch
12 Sep 2020
Journal of Investigative Dermatology | VOL. 141

Sample effficient deep reinforcement learning for control

-

15 Dec 2019
15 Dec 2019

Accelerating Deep Learning Tasks with Optimized GPU-assisted Image Decoding
Lipeng Wang ... Shengen Yan
-
Lipeng Wang, et. al.Lipeng Wang ... Shengen Yan
01 Dec 2020
01 Dec 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Automatic data acquisition for deep learning

Abstract

Talk to us

Similar Papers

More From: Proceedings of the VLDB Endowment