In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.
In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.
https://en.wikipedia.org/wiki/Lasso_(statistics)#Geometric_i...
The derivative of that loss with respect to x would be J'(x) + beta * sgn(x). So using some variant of SGD (which is what basically all neural network training does these days) we would essentially update x as: x = x - alpha * (J'(x) + beta). (The specifics depend on the algorithm, but it doesn't change the result).
So for x to end up as _exactly_ 0, we have to be extremely lucky, which in practice I have never observed. Using L1 regularization definitely leads to small weights, but not to ones that are _exactly_ 0.
You can prove that L1 regularization is equivalent to taking the optimal unregularized parameters, setting to parameters below a threshold to 0 (the threshold depends on the regularization parameter), and penalizing the other parameters.
2) When you use actual numerical optimization techniques with floating point arithmetic, you don't find an exact minimum (global or local). And you don't get exact zeros.
Have you tried this on a real problem? I wouldn't consider MNIST a real problem, but even there you will not get _exact_ zeros. Try it.