Interpreting neural networks through the polytope lens (2022)
lesswrong.com
lesswrong.com
Of course, the relu unit is also a passing on information when the result is on one side of the plane, making this a spline.
As others have said... Can we learn the separating planes without the backward gradient propagation? I don't know but seeing it in this new way may help.
https://www.nature.com/articles/d41586-024-00288-1
I would love to see a cross section of these two ideas....
I have been surprised that in the past few weeks, I have seen several posts on HN where, while separate, unrelated posts here - there have been related characteristics and if you look at them for a sec, you can see how having AIs GPT both studies/papers - immediate connections worth looking at further are revealed.
If even for the sake of just a more informed tapestry of knowledge in a particular area...
Its really enjoyable reading and TIL'ing.
Normalization removes this problem. Magnitude information can still be encoded separately in a log form so differentiation can still happen when scale matters, but scaling doesn't have much impact by default (small initial weights following magnitude element).
[0]: https://proceedings.neurips.cc/paper/2018/file/22b1f2e098316... pro
(Edited for some clarity)