The boundary of neural network trainability is fractal
arxiv.org
arxiv.org
I see you
> Have you ever done a dense grid search over neural network hyperparameters? Like a really dense grid search? It looks like this (!!). Blueish colors correspond to hyperparameters for which training converges, redish colors to hyperparameters for which training diverges.
So what this is saying, for certain ill-chosen learning weights, model convergence is for lack of a better word, chaotic and unstable.
Training consists of 500 (sometimes 1000) iterations of full batch steepest gradient descent. Training is performed for a 2d grid of η0 and η1 hyperparameter values, with all other hyperparameters held fixed (including network initialization and training data).
This is really fun to see. I love toy experiments like this. I see that each plot is always using the same initialization of weights, which presumably makes it possible to have more smoothness between each pixel. I also would guess it's using the same random seed for training (shuffling data). I'd be curious to know what the plots would look like with a different randomness/shuffling of each pixel's dataset. I'd guess for the high learning rates it would be too noisy, but you might see fractal behavior at more typical and practical learning rates. You could also do the same with the random initialization of each dataset. This would get at if the chaotic boundary also exists in more practical use cases.
one would have to do many runs for each point in the grid and average them or something.
but i didn't read the paper so maybe they did.
We also know the brain, cortex esp, is highly recurrent, so it should be primed for creating fractals and chaotic mixing.
So maybe the hidden structure is the set of neural hyperparams needed to put a given cluster of neurons into fractal/chaotic oscillations like this. Seems potentially more useful too.. way more information content than a configuration that yields a fast convergence to a fixed point.
Perhaps this is what learning deep NNs is doing: producing conditions where the substrate is at the tipping point, to get to a high-information generation condition, and then shaping this to fit the target system as well as it can with so many free parameters.
That suggests that using iterative generators that are somehow closer to the dynamics of real neurons would be more efficient for AI: it'd be easier to drive them to similar feedback conditions and patterns
Like matching resonators in any physical system
Well, they both also have similar fractal dimensions. Mandlebrot's Hausdorff dimension is 2, Julia is 1-2.
I won't argue it here but just suggest that this is an important complexity relationship and that the neural net being fit may also have similar fractal complexity, and that the distinction between param and hyper-param in this sense may be somewhat a red-herring
It’s kind of the opposite, no? The Mandelbrot set is the set of points where the Julia set for that point includes that point.
https://acko.net/blog/how-to-fold-a-julia-fractal/
Cheers!
I’m sure from your comment you are aware of the distinction, but it is an interesting concept for people to keep in mind.
2) not sure what you mean by "operate in discrete space"
I'd emphasize the potential similarity to biological recurrence. Deep ANNs don't need to have this explicitly (tho e.g. LSTM has explicit recurrence), but it is known that recurrent NNs can be emulated by unrolling, in a process similar to function currying. In this mode, a learned network would learn to recognize certain inputs and carry them across to other parts of the network that can be copies of the originator, thus achieving functional equivalence to self feedback, or neighbor feedback. It takes a lot of layers and nodes in theory, but ofc modern nets are getting very big.
Text is infinitely complex. Written text is only somewhat less so. The text people choose to write is full of complex entropy.
The language patterns we use to read text, on the other hand, are much more simple. The most complicated part is ambiguity, and we don't resolve that with language. Instead, we resolve ambiguity with context.
It's a similar problem with "AI". Artificial Intelligence does not exist, yet we put that name on all kinds of things. This causes real problems, too: by labeling an LLM "AI", we anthropomorphize it. From then on, the entire narrative is misleading.
An LLM is a model, not an actor. As soon as we call it "AI", that distinction gets muddled, and the whole narrative follows.
No, it isn’t.
> An LLM does not model language. The name is misleading, and should be changed to Large Text Model.
What, because it isn’t processing speech?
Presumably you are aware of models that model recordings of speech, and therefore this isn’t what you mean.
In that case, it seems to me like you are probably jerrymandering some concepts?
An LLM is a model of written text. It doesn't know anything about language rules. In fact, it doesn't follow any rules whatsoever. It only follows the model, which tells you what text is most likely to come next.
A program that has a good rate at distinguishing whether a string is human written, would be substantially longer than one that recognizes the language of all strings over a particular alphabet.
If you want to generate strings instead of recognizing them, a program that enumerates all possible strings over a given alphabet, can also be pretty short.
Not sure what you mean by complexity.
I don’t know what you mean by “it doesn’t follow any rules at all”.
When you train an LLM, you don't write any language rules. Instead, you provide examples of written text. This approach is implicit: syntax is not known ahead of time. The core feature is that there is no distinction between what is correct and what is not correct. The LLM is liberated from the syntax rules of language. The benefit is that it can work with ambiguity. The limitation is that it can't decide what interpretation of that ambiguity is correct. It can only guess what interpretation is most likely, based on the text it was trained on.
... Sorry I couldn't help myself.
It relates to fractal (non-integer) dimensions, which was first described by Mandelbrot in a paper about self similarity.
Here is a paper that covers some of that.
https://www.minvydasragulskis.com/sites/default/files/public...
In Newton's fractal, no matter how small a circle you can draw, your circle will either contain one root or all the roots.
The basins that contain one root are open sets that share a boundary set.
Even if you could have perfect information and precision this property holds. This means any change in initial conditions that crosses a boundary will be indeterminate.
There is another feature called riddled basins, where every point is arbitrarily close to other basins. This is another situation where even with perfect information and unlimited precision a perturbations would be indeterminate.
A positive Laponov exponent which isn't sufficient to prove chaos, but is always positive in the presence of chaos may even be 0 or negative in the above situations.
Take the typical predator prey model and add fear and refuge and you hit the riddled basins.
Stack four reflective balls in a pyramid and shine different color lights in two sides and you will see the Wada property.
Neither of those problems are addressable with the assumption of deterministic effects with finite precision.
Reading this gave me goosebumps
Yes, I also think this is strange. In regular fractals the x and y coordinates have the same units (roughly speaking), but here this is not the case, so I wonder how they determine the relative scale.
If you envision a given architecture as a class of (higher-order) function, the inputs would be the parameters, and the constants would be the hyperparameters. Varying the constants moves to a different function in the class, or, varying the hyperparameters gives a different model with the same architecture (even with the same data).
https://en.wikipedia.org/wiki/Julia_set
or
The parameters that you tweak to control model learning have a self-similar property where as you zoom if you see more and more complexity. Its the very definition of local maxima all over the place.
And indeed, the AI response indicated that the boundary between convergence and divergence is not well defined, has many local maxima and minima, and could be quote: "fractal or chaotic, with small changes in hyperparameters leading to drastically different outcomes."
If you watch this video[0], you'll see in the first frame that there is a clear boundary between learning rates that converge or not. Ignoring this paper for a moment, what if we zoom in really really close to that boundary? There are two possibilities, either (1) the boundary is perfectly sharp no matter how closely we inspect it, or (2) it is a little bit fuzzy. Of those two possibilities, the perfectly sharp boundary would be more surprising.
It's not only the boundary that is fractal.
We'll soon see that learning on one dataset (area of fractal) with enough data will generalize to other seemingly unrelated datasets.
There is evidence that the structure neural networks are learning to approximate in a generative fractal of sorts.
Finally, we'll need to adapt gradient descent to operate at move between different scales