It sounds like there is a growing bag of tricks neural network researchers are discovering to make training practical and stable for large data sets.
One example would be using relu activation - whenever I play with it in a simple tutorial like this one, training seems to explode and fail much more frequently, so I'm guessing either I'm missing another step people use, or there are some extra constraints on initial conditions?
Using a Gaussian for activation in my tutorials has tended to be more stable and converge much faster, but I assume there is a huge downside lurking somewhere to having a non-monotonically increasing function?
What are the tricks of the trade that a weekend warrior should investigate?