And his visualization of constrained optimization is astonishing https://explained.ai/regularization/index.html (I struggled for a long time to get the right intuition of a Lagrangian)
Thanks! Took me a year to discover the key nut there. L1 vs L2 regularization is not well described I found so I went nuts trying to nail it down.
If you're interested, in my thesis I induced l1-regularized decision trees through a boosting style approach. Adding an l1 term and maximizing the gradient led to sparse tree.