The generally accepted definition in the community is Tom Mitchell's:
> A computer program is said to learn from experience E with respect to some task T and performance measure P, if its performance at task T, as measured by P, improves with experience E.
Statistical estimation methods are one way to achieve this, but not the only way, especially when an exact function can be learned, i.e. a typical layer 2 switch is a learning device. You don't program in the mapping from connected device MAC addresses to switch port, the switch itself learns this from receiving a message from a MAC on a port and then records the mapping. That is a very simple form of non-statistical machine learning.
I'm not really sure how you can start here but then say regression is not a form of machine learning. "Regression" is a pretty broad class of techniques that just means using labeled examples to estimate a function with a continuous response, typically contrasted with "classification," where the response is discrete. The method you use to do the function approximation may or may not be statistical. A genetic algorithm is not, for instance. I'm not sure least squares, which is what Legendre invented, should really be considered statistical, either. The original purpose was for approximating the solution to an overdetermined system of equations, before statistics was even formalized. It certainly became a preferred method in mathematical statistics, but mathematical statistics wasn't even formalized until later. It didn't start being called "regression" until Galton published his paper showing children of exceptionally tall or short people tended to regress to the mean, which was 80 years later. But you're performing regression analysis whether you use the normal equations, LU decomposition, QR decomposition, a genetic algorithm, gradient descent, stochastic gradient descent, stochastic gradient descent with dropout. Doesn't matter. As long as you're doing function approximation with a continuous response, it's still regression. Whether or not it can also be considered "machine" learning just depends on whether you're doing it by hand or via machine.
Though sure, typically people tend to imagine the more exotic and newer techniques that scale to very large data sets and reduce overfitting and deal with noise automatically and involve hyperparameters, i.e. not least squares.