Proceeds to use (1024^2 * 2 + 1024) parameters in the neural network.
Proceeds to use (1024^2 * 2 + 1024) parameters in the neural network.
As the author is using RELU, he will have a decent number of neurons 'die'. So some 'over-provisioning' is not a bad idea in theory. Also, if my keras isn't so rusty, I think the author is using less parameters than you are stating.
Still a more reasonable dropout rate and maybe some regularization/batch normalization might help, but I would say not over fitting on only 700 samples is a hard task, even with a network much smaller than that.
I would be pretty shocked if this neural net wasn't over fit
[1]-https://towardsdatascience.com/pruning-deep-neural-network-5...
I just thought I'd highlight a bit of funniness.
/s
/s for this post, I mean. I have had this very suggestion made to me non-sarcastically under similar circumstances.
after many similar comments I plan on implementing a simpler model and seeing how it compares