The famous ReLU non-linearity is just that - two linear functions joined.
Layers reduce resource requirements and make some patterns easier or even practical to find, but any ANN that is a FNN supervised learning could be represented as a parametric linear regression.
Unsupervised learning, that tends to use clustering is harder to visualize but is the same thing.
You still have ANNs, which have binary output, which can be viewed through the lens of deciders. They have to have unique successor and predecessor functions.
Really this is just set shattering that relates to a finite VC dimensionality being required for something to be PAC learnable.
But the title of this is confusing the map for the territory. It isn't that 'Everything is a linear model' but that linear models are the preferred, most practical form.
The efforts to leverage spikey neutral networks, which is a more realistic model of cortical neurons, and which have continuous output (or more correctly the computable reals) tend to run into problems like riddled basins.
https://arxiv.org/abs/1711.02160
Obviously setting rectified linear unit at 0 = 1 resolves to differentiation problem, but many functions may not be so simple
Perhaps a useful lens is how TSP with a discreet Euclidean metric is in NP-complete while the continuous version is in NP-hard.
But it isn't that everything is linearizable, but rather that linearized problems tend to be the most practical.
Consider Predator Pray with fear and refuge, which is indeterminate, and not due to a lack of precision but a topological feature where ≥3 open sets share the same boundary set.
https://www.sciencedirect.com/science/article/abs/pii/S09600...
General relativity, with 3 spacial and one temporal dimension is another. One lens to consider this is that rotations are hyperbolic due to the lack of independence from the time dimension.
Quantum mechanics would have been much more difficult if it didn't have two exit basins. Which is similar to ANNs and linear regressions being binary output.
(Some exceptions will orthogonal dimensions like EM)
A neural network can be pedantically referred to as a linear model of the form y = a + b*neural_network, for example. Here, y is a linear model (even though neural_network isn't).