Everything is a linear model
danielroelfs.com
danielroelfs.com
My series of lm/glm/gam/gamm revelations was:
1. All t-tests and ANOVA flavors are just linear models
2. Linear models are just a special case of generalized linear models (GLMs), which can deal with binary or count data too
3. All linear models and GLMs are just special cases of generalized linear mixed models, which can deal with repeated measures, grouped data, and other non-iid clustering
4. Linearity is usually a bad assumption, which can easily be overcome via splines
5. Estimating a spline smoothing penalty is the same thing as estimating the variance for a random effect in a a mixed model, so #3 and #4 can be combined for free
And then you end up with a generalized additive mixed model (GAMM), which can model smooth nonlinear functions of many variables, smooth interaction surfaces between variables (e.g. latitude/longitude), handle repeated measurements and grouping, and deals with many types of outcomes, including continuous, binary yes/no, count, ordinal categories, or survival time measurements.
All while yielding statistically valid confidence intervals, and typically only taking a few minutes of CPU time even on datasets with hundreds of thousands / millions of datapoints.
Is there some go-to practice material I could look at?
Splines I haven't touched since numerical computing exercises in school.
I have also heard great things about Frank Harrell's Regression Modeling Strategies which uses a slightly different approach (still spline-based though), but I haven't read it. His other writing is fantastic though.
Very very good indeed, with the exception that he basically ignores compute time and efficiency. I learned a bunch, but applying his approach to the kinds of datasets I deal with (much larger and with pretty strict compute budgets) was very difficult.
You might also enjoy reading the "Neural Additive Model" paper from Hinton's lab, which is basically GAMs using a separate DNN as a "spline basis" for each input variable.
You aren't guaranteed that your equilibria behave the same in the linearization of a nonlinear system if your Jacobian has any eigenvalues with real part of 0.
https://en.wikipedia.org/wiki/Hartman%E2%80%93Grobman_theore...
Fun times.
Don't have the code or writeup any more.
y = b1*x1 + b2*x2 + ... + bp*xp
which is in the heart of the model (perhaps more specific description would be linear combination). Whatever you call it, it's just some quantity y, that is constructed via an additive process from components x1...xp, and each component is multiplied by some constant (the coefficient of the linear combination).This linear combination "core" of the model can be used directly (in which case it is geometrically a hyperplane), but you could also use non-linear transformation of the inputs or the outputs.
E.g. x1 could be a spline basis function, or log-transformed data, or anything.
Similarly, the output y can be passed though some non-linearity to force a particular output (like in logistic regression). Hence the statement is not vacuous
Maybe what OP meant is that particular non linearities (splines) allow to keep some goodies from the linear model in non-linear settings. Still it's not clear for me in which settings exactly such models are good enough
within set boundaries.
So far I have been successfully using GLMMs where appropriate, but then jumped to implementing completely arbitrary models by fitting them to the data (plus bootstrapping).
But that takes ages for most problems as you say
and everything turns around the same principles. For example dynamical models and PID controls.
yet solving a banana, is the only thing we really know how to do. So we end up fitting everything in our banana models.
every smooth system, sure, but even continuity is no guarantee of locally linearity.
I'm skeptical that discontinuities can exist because, if they did, they'd serve as infinitely powerful microscopes. If there's a discontinuity in nature, it must exist at absolute zero. I don't have a similarly good argument for continuous first derivatives but I do think it's interesting that there are no examples AFAIK.
Study it in school?
There are a lot of linear problems that I'm interested in.
Do you use a computer algebra system?
Do you have a few books?
We just break the problem in small bananas because that's what we can solve, and then solve for those lil'bananas and call it done.
At the time I was mad at engineers for being non-scientific, but after a few years I understood the deep wisdom in that. Nonlinear materials exist, and materials we use have nonlinear ranges. We just don't build things from those, because the math is too unwieldy. (except in very very specific edge cases where we spend a lot of money building a very limited thing)
So I left the "physics" courses and went the "math" courses... only to learn 3 years later how to prove that this kind of approximation is mathematically sound indeed :-D
I'm linking to my fork of it because I've added some fixes and filled in some of the missing parts (e.g. Welch's t-test).
It really makes me frustrated about the ways I was introduced to statistics: brute force memorization of seeming arbitrary formulas.
Not the case with non-linear models. We need to throw computers at them.
If the model is y = f(x) + e, where for example f is a neural net, there is not much we can say about the “quality” of parameters of f. That is, we can’t attach confidence interval. We will have to resort to expensive simulations and for most practical applications this is not workable (size of dataset).
What you say is useful in a different context. Local linearization is a very powerful idea.
estimating the Area Under the Curve metric (AUC) is equivalent to the Wilcoxon-Mann-Whitney test!
https://rmets.onlinelibrary.wiley.com/doi/abs/10.1256/003590...
A neural network can be pedantically referred to as a linear model of the form y = a + b*neural_network, for example. Here, y is a linear model (even though neural_network isn't).
The famous ReLU non-linearity is just that - two linear functions joined.
Layers reduce resource requirements and make some patterns easier or even practical to find, but any ANN that is a FNN supervised learning could be represented as a parametric linear regression.
Unsupervised learning, that tends to use clustering is harder to visualize but is the same thing.
You still have ANNs, which have binary output, which can be viewed through the lens of deciders. They have to have unique successor and predecessor functions.
Really this is just set shattering that relates to a finite VC dimensionality being required for something to be PAC learnable.
But the title of this is confusing the map for the territory. It isn't that 'Everything is a linear model' but that linear models are the preferred, most practical form.
The efforts to leverage spikey neutral networks, which is a more realistic model of cortical neurons, and which have continuous output (or more correctly the computable reals) tend to run into problems like riddled basins.
https://arxiv.org/abs/1711.02160
Obviously setting rectified linear unit at 0 = 1 resolves to differentiation problem, but many functions may not be so simple
Perhaps a useful lens is how TSP with a discreet Euclidean metric is in NP-complete while the continuous version is in NP-hard.
But it isn't that everything is linearizable, but rather that linearized problems tend to be the most practical.
Consider Predator Pray with fear and refuge, which is indeterminate, and not due to a lack of precision but a topological feature where ≥3 open sets share the same boundary set.
https://www.sciencedirect.com/science/article/abs/pii/S09600...
General relativity, with 3 spacial and one temporal dimension is another. One lens to consider this is that rotations are hyperbolic due to the lack of independence from the time dimension.
Quantum mechanics would have been much more difficult if it didn't have two exit basins. Which is similar to ANNs and linear regressions being binary output.
(Some exceptions will orthogonal dimensions like EM)
You would have fitted right in.
You might not even need anything else…
https://www.reddit.com/media?url=https%3A%2F%2Fi.redd.it%2F4...