Say we have a predictor matrix X with 2 predictors. We fit a model using a penalized linear regression (say LASSO) adding to our predictor matrix an interaction terms, arbitrary polynomial and logarithmic transformations of each X, and interactions between the transformations of the Xes. Ideally we motivate this because of some case knowledge about relevant nonlinear transformations of the predictors. Or maybe second best we use the kernel trick to run a kernelized regression that uses an infinite dimensional prediction of all possible transformations of the predictors. But realistically, we toss some shit in the model and run it.
The LASSO spits out X1, X2^2, and X1^3 * log(X2) as being the cross-validation selected non-zero parameters.
What real world scenario could possibly generate a causal process that is linear in X1 (say income), quadratic in X2 (say age, which often displays quadratic forms in regressions), but also predicted by a bizarre non-linear interaction of nonsense transformations?
What a practitioner would probably do is fit the model. In a lot of ML contexts, interpretation would lead to the practitioner saying "Well, okay, ML sometimes produces nonsense models, but you can't argue with the predictive results". Or maybe the practitioner is more sensitive to interpretability and instead takes another tack. Maybe the practitioner might say "clearly this interaction is nonsense, but there must be some interaction, I'll re-run with a linear interaction". Or else they'd re-run the LASSO with conditions about not including nonlinear terms without including the lower dimensional terms. Or else they'd run a grouped LASSO and make up some justification for the groups. All of these reveal that most ML practitioners are basically just doing alchemy.
And this is talking about what amounts to a minor version increment of linear regression, so probably the simplest possible technique we'd still call part of ML.
If so, then Machine Learning is a part of modeling AI. Regardless of how they are taught in terms of University lectures.
In fact, in real data, these assumptions are almost always violated. The Gaussian assumption doesn't matter at all, but to address the i.i.d. assumption: Almost all real data exhibits residual heteroskedasticity and almost all real data has observable clustering. Which is why almost no one uses OLS with classical errors. We have estimators to allow errors to be heteroskedasticity-consistent (the default in STATA and easily estimated in R e.g. by estimatr, clubSandwich, etc) or cluster-robust or both. By definition these cases have non-i.i.d. errors and there's no reason linear regression can't be used with them.
We also don't need to make assumptions, these can be interrogated. Most regression relies on using the residual matrix as sample plug-ins for the underlying error matrix, so there's a wide assortment of diagnostic techniques to check for the presence or absence of those assumptions.
Insofar as "machine learning" has any meaning -- which is to say, insofar as it is different than "statistics", the difference is purportedly that it focuses on minimizing out of sample prediction error rather than estimating population parameters, and typically this is motivated as an overfitting problem.
We use OLS because OLS is BLUE under the Gauss-Markov conditions. In ML we rarely care about "U" (unbiasedness) because we frequently prefer to make a bias-variance tradeoff if we're aiming to minimize out of sample error. When linear regression is used in an ML context it is typically penalized linear regression (i.e. ridge / LASSO). Of course it's also the case that the bulk of sexy ML results come out of non-linear estimators, and absent a need to characterize population parameters there's no real reason to care about interpretability so really we don't care about the "L" either.
I would say the grandparent is closer to right. Often in ML there is a view that we throw a bunch of processes at data, pick the thing that works best, don't care why it works at all, and then run with it. To the extent there's a protection against fishing expeditions, it's in the training/test separation or cross-validation or both.
Most of the time when someone talks about "regression theory", they're used 30 or 40 year old results. For an updated look, check out "Foundations of Agnostic Regression" (Aronow and Miller, both Yale Political Scientists) which is coming out some time in 2019. They've had a pre-print around for a while and if you're interested I'm sure you could get one.
- IID: "independent and identically distributed", https://en.wikipedia.org/wiki/Independent_and_identically_di...
- OLS: "ordinary least squares", https://en.wikipedia.org/wiki/Ordinary_least_squares (I think)
The grandparent to your content was raising an objection that, actually, linear regression, a very old technique which in plain English means fitting a straight line to a scatterplot (but in any number of dimensions), has a great deal of theory around it. The simplest form of solving a linear regression is by "ordinary least squares" (minimizing the sum of squared deviations from the fit line): OLS.
The grandparent was correct that especially in the mid-20th century to late 20th century, a lot of people did work on the conditions under which OLS works. What "works" means in a statistical sense is that it's efficient (has low uncertainty about the correct estimate), unbiased (on average gets the right answer), consistent (as you have more and more data gets closer to the right answer). Under a set of fairly impossible conditions about the real world data generated process, OLS is "BLUE" (the best linear unbiased estimator). Best here refers to efficiency, and unbiasedness I've already explained. OLS divides the data into structural elements (things that can be explained by the predictors you put into the estimator) and stochastic elements (the noise left over -- the deviations from the line). If we specify the correct model, the stochastic elements are the underlying stochasticity in the universe. If we specify an incorrect model, some of our omitted structural elements get put into the estimation of the noise.
The grandparent noted that two assumptions made in linear regression are that the underlying stochastic disturbances in the data are i.i.d. (each is a random draw from the same distribution) and Gaussian (form a normal / bell curve). These are not assumptions, these are conditions under which OLS is BLUE. The latter is not a necessary condition at all, the distribution can take any form. The former is the most succinct way to express one of the conditions.
My comment was to raise that actually when we use linear regression in the world, we rarely use classical OLS. In the real world underlying disturbances differ between observations. Imagine if I am running a regression on cross-country data, but while my US data is very precisely measured thanks to the widespread availability of polling firms, my Mexican data involves census enumerators going to rural villages. We might imagine that all of my US data is more precisely measured than all of my Mexican data, so we would expect the underlying stochasticity to differ between country. This is called clustering. Also, because we almost certainly do not have the correct model for the data (say log-dollars income predicts the result, not dollars income, but I put in dollars income), it can be the case that observations with higher values of our predicted Y also have more uncertainty. This is called heteroskedasticity. But the good news is we have answers to both, we just don't use OLS, we use more modern estimators. Yay!
In general the world has moved away from rigorously teaching the conditions under which OLS and works and toward teaching more flexible estimators that work under less restrictive conditions. And in general in ML, people aren't using anything that looks anything like OLS, because ML has specific goals OLS is inappropriate for -- namely minimizing overfitting and out-of-sample error, where OLS is designed to maximize the precision of estimates of the slope parameters (how a given predictor affects the outcome). So all the work in OLS theory doesn't really translate to a machine learning setting, where many methods have no theory at all.
Hope this is a plainer English version.