HNHacker News
TopNewBestAskShowJobs

rvarma

74 karma · joined October 19, 2017

submissionscomments
rvarma··on Batch Normalization for deep networks
Ah yeah, my bad, I should've instead shown the activations after the first layer, since "activation 0" is just the distribution of the random data I started with
rvarma··on Batch Normalization for deep networks
I actually think the idea of using leaky ReLUs is interesting, because it'll still provide a small gradient when x < 0, which perhaps may slightly alleviate the vanishing gradients issue
rvarma··on Batch Normalization for deep networks
Thanks for your comment!

Regarding point 1, I stored both the activations before batchnorm and after since I needed them during the backwards pass. i.e. I stored h_out before and after these operations:

h_out = (h_out - np.mean(h_out, axis = 0)) / np.std(h_out, axis = 0) h_out = gamma * h_out + beta

Regarding point 2, I do realize now that fixing a bad init via Xavier/he initialization and using batch norm fix slightly different problems - if I were to rewrite this post I probably wouldn't talk about initialization at all, or at least mention the Xavier/He initialization.

rvarma··on Language Models, Word2Vec, and Efficient Softmax Approximations
Thanks for the link! It actually provides a really clear and intuitive explanation of the notion of similarity.