A Neural Network in 11 Lines of Python (2015)
iamtrask.github.io
iamtrask.github.io
What I'm talking about is the size that is required so that the neural net can learn it. This may be different.
It sounds like there is a growing bag of tricks neural network researchers are discovering to make training practical and stable for large data sets.
One example would be using relu activation - whenever I play with it in a simple tutorial like this one, training seems to explode and fail much more frequently, so I'm guessing either I'm missing another step people use, or there are some extra constraints on initial conditions?
Using a Gaussian for activation in my tutorials has tended to be more stable and converge much faster, but I assume there is a huge downside lurking somewhere to having a non-monotonically increasing function?
What are the tricks of the trade that a weekend warrior should investigate?
Yes, I'm taking the specialization and having a blast with it. :-)
[1] https://www.coursera.org/learn/neural-networks-deep-learning
The assignments are excellent and will let you implement a deephish network from practically scratch, including backprop, optimizers, tuning etc.
In any case, reposting is allowed, for several good reasons, which have been discussed in the past.