Ah yeah, my bad, I should've instead shown the activations after the first layer, since "activation 0" is just the distribution of the random data I started with
74 karma · joined October 19, 2017
Regarding point 1, I stored both the activations before batchnorm and after since I needed them during the backwards pass. i.e. I stored h_out before and after these operations:
h_out = (h_out - np.mean(h_out, axis = 0)) / np.std(h_out, axis = 0) h_out = gamma * h_out + beta
Regarding point 2, I do realize now that fixing a bad init via Xavier/he initialization and using batch norm fix slightly different problems - if I were to rewrite this post I probably wouldn't talk about initialization at all, or at least mention the Xavier/He initialization.