Could anyone help me on what basis the number of hidden layers are chosen for deep neural network.
Show older comments
Let me explain in brief.
I have generated the code for deep neural network for regression purpose using numerical data to predict the formation of clusters.
when I run the code, for four hidden layers i can get the lowest value of mean square error as compared to 2 hidden layers,3 hidden layers,5 hidden layers, and 6 hidden layers.
So,I can say four hidden layers are optimal in my case.
But I would like to know is there any other reason other the mean square error to justify why four hidden layers are optimal.
Also let me know, for an image based on pixel, I can find low level features, high level features and so on.
But for numerical data what represent low level and high level features.
Could anyone please clarify me.
Answers (1)
But I would like to know is there any other reason other the mean square error to justify why four hidden layers are optimal.
No, if you change the loss function or any other thing about your network architecture (e.g., number of neurons per layer), you could very well find you get a different optimal number of layers.
But for numerical data what represent low level and high level features.
In general, you won't know that in advance. The main purpose of a neural network is for the network to learn the relevant features on its own are during training.
15 Comments
jaah navi
on 4 Jul 2021
jaah navi
on 4 Jul 2021
Matt J
on 4 Jul 2021
Also, the code...
jaah navi's comment moved here:
Code:
XTrain
YTrain
inputSize = 2;
numHiddenUnits1 = 20;
numHiddenUnits2 = 20;
numClasses = 1;
layers = [ ...
sequenceInputLayer(2)
fullyConnectedLayer(20)
reluLayer
fullyConnectedLayer(20)
reluLayer
fullyConnectedLayer(1)
regressionLayer]
maxEpochs = 50;
miniBatchSize = 50;
options = trainingOptions('adam', ...
'ExecutionEnvironment','cpu', ...
'MaxEpochs',maxEpochs, ...
'MiniBatchSize',miniBatchSize, ...
'GradientThreshold',1, ...
'InitialLearnRate',0.001,...
'Verbose',false, ...
'Plots','training-progress');
net = trainNetwork(XTrain,YTrain,layers,options);
Matt J
on 4 Jul 2021
Yes, but also for the 30 unit case, and also it would be good to see the training curves, since we don't have your data to train with.
jaah navi
on 4 Jul 2021
I don't really see a difference in the final RMSEs that I can be certain is signficant. They could just differ by random fluctuations inherent to stochastic gradient descent. You might try reducing the learning rate to see if you can get better descent, however.
jaah navi
on 4 Jul 2021
jaah navi
on 4 Jul 2021
Categories
Find more on Deep Learning Toolbox in Help Center and File Exchange
Products
Community Treasure Hunt
Find the treasures in MATLAB Central and discover how the community can help you!
Start Hunting!