Neural architecture search via standard machine learning methodologies

Giorgia Franchini,Luca Zanni,Valeria Ruggiero,Federica Porta

doi:10.3934/mine.2023012

Abstract

<abstract><p>In the context of deep learning, the more expensive computational phase is the full training of the learning methodology. Indeed, its effectiveness depends on the choice of proper values for the so-called hyperparameters, namely the parameters that are not trained during the learning process, and such a selection typically requires an extensive numerical investigation with the execution of a significant number of experimental trials. The aim of the paper is to investigate how to choose the hyperparameters related to both the architecture of a Convolutional Neural Network (CNN), such as the number of filters and the kernel size at each convolutional layer, and the optimisation algorithm employed to train the CNN itself, such as the steplength, the mini-batch size and the potential adoption of variance reduction techniques. The main contribution of the paper consists in introducing an automatic Machine Learning technique to set these hyperparameters in such a way that a measure of the CNN performance can be optimised. In particular, given a set of values for the hyperparameters, we propose a low-cost strategy to predict the performance of the corresponding CNN, based on its behavior after only few steps of the training process. To achieve this goal, we generate a dataset whose input samples are provided by a limited number of hyperparameter configurations together with the corresponding CNN measures of performance obtained with only few steps of the CNN training process, while the label of each input sample is the performance corresponding to a complete training of the CNN. Such dataset is used as training set for a Support Vector Machines for Regression and/or Random Forest techniques to predict the performance of the considered learning methodology, given its performance at the initial iterations of its learning process. Furthermore, by a probabilistic exploration of the hyperparameter space, we are able to find, at a quite low cost, the setting of a CNN hyperparameters which provides the optimal performance. The results of an extensive numerical experimentation, carried out on CNNs, together with the use of our performance predictor with NAS-Bench-101, highlight how the proposed methodology for the hyperparameter setting appears very promising.</p></abstract>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Mathematics in Engineering	Publication Date: Jan 1, 2022
Citations: 13	License type: cc-by

R Discovery Prime

R Discovery Prime

Neural architecture search via standard machine learning methodologies

Abstract

Talk to us

Similar Papers

More From: Mathematics in Engineering

Lead the way for us

Similar Papers

Deep learning-based roadway crack classification using laser-scanned range images: A comparative study on hyperparameter selection
Shanglian Zhou ... Wei Song
Automation in Construction | VOL. 114
Shanglian Zhou, et. al.Shanglian Zhou ... Wei Song
18 Mar 2020
Automation in Construction | VOL. 114

Sensitivity Analysis of the Hyperparameters of CNN for Precipitation Downscaling
Takeyoshi Nagasato ... Masato Kiyama
-
Takeyoshi Nagasato, et. al.Takeyoshi Nagasato ... Masato Kiyama
03 Mar 2021
03 Mar 2021

The structural tuning of the convolutional neural network forspeaker identification in mel frequency cepstrumcoefficients space
Anastasiia D Matychenko ... Marina V Polyakova
Herald of Advanced Information Technology | VOL. 6
Anastasiia D Matychenko, et. al.Anastasiia D Matychenko ... Marina V Polyakova
03 Jul 2023
Herald of Advanced Information Technology | VOL. 6

Predicting How CNN Training Time Changes on Various Mini-Batch Sizes by Considering Convolution Algorithms and Non-GPU Time
Peter Bryzgalov ... Toshiyuki Maeda
-
Peter Bryzgalov, et. al.Peter Bryzgalov ... Toshiyuki Maeda
21 Jun 2021
21 Jun 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Neural architecture search via standard machine learning methodologies

Abstract

Talk to us

Similar Papers

More From: Mathematics in Engineering