Pre-Trained Large Language Models Use Fourier Features for Addition (2024)
arxiv.org
arxiv.org
Interpreting Modular Addition in MLPs https://www.lesswrong.com/posts/cbDEjnRheYn38Dpc5/interpreti...
Paper Replication Walkthrough: Reverse-Engineering Modular Addition https://www.neelnanda.io/mechanistic-interpretability/modula...
So saying it's not the simplest is an understatement by far, it's doing millions or billions of times as much math as a calculator would. Ask an LLM to generate a program to do math, rather than doing the math itself.
1. Well somebody has got to do the work and we can't all just go around assuming stuff, even if we're pretty confident. The confirmation helps and is beneficial to the community.
2. It's obvious post hoc and you have gaslit yourself into thinking that you already knew it because you kinda knew it at a high level and you only read the result at a high level too so you entirely miss all the actual details and all the context (especially since there is so much context that never makes it into a paper[0])
3. (bonus) It addresses the same thing that someone else addressed but from a different approach and the new approach can provide additional insight.
Either way, it is beneficial to the community. Sure, nothing is groundbreaking but that's how science is. 99% incremental steps. And hey, these days most ML papers are just an active demonstration of how well you can brute force search optimal hyperparameters (i.e. how much compute you can afford). I see far fewer sufficiently isolating variables and actually provide strong evidence of the things claimed as novel, not recognizing that benchmark results are far from sufficient. But I blame reviewers for that, but also see rant in [0][0] I think the hardest thing about beginning a PhD is that you're reading a bunch of papers going like "why the fuck are they doing this?" and the problem is that you don't have enough breadth or depth to get it. You don't understand the decades long conversation of how we got here and what problems were being addressed along the way. To be fair, a lot of this is never stated explicitly and so you annoyingly have to piece everything together by reading a few hundred papers. But also, good luck providing all that context within page limits and besides, papers are written for /peers/, by which I mean niche peers, not domain peers. Ain't nobody got time to write textbooks, because you're just trying to publish so you don't perish and you're already exhausted from all the grant writing, rebuttals, bureaucratic work, and all that fun jazz.
If that's our metric then most humans haven't truly learned addition either
For any neural network, the standard you can expect for any learned skill is closer to a human learning that skill than to a computer programmed to do that thing. There will be occasional mistakes
I'm sure the llm could formally do it too
#1: For small numbers, the answer is "directly" available to us without consciously applying any steps
#2: For larger numbers, we consciously apply some formal method to break the problem down into smaller steps of type #1. Like column addition, adding digit-by-digit and carrying the ones
Despite seeming direct, I'd argue that if we could see at a low-level what our neurons were actually doing for #1, like to get 3 + 5, it would likely also look roundabout. Might even be a similar process as with LLMs, approximating magnitude (~7-9) then snapping to parity (even, since 3 and 5 are odd).
LLMs should be capable of #2, including choosing an appropriate method, with chain-of-thought reasoning. But in addition to that, and I think what bongodongobob is getting at, is that LLMs appear to have a more robust #1 than us - being able to accurately add far larger numbers whereas we'd normally fall back to a step-by-step method after one or two digits.
It is also an extremely neat piece of the real world, but I’m hesitant to guess your background and offer an explanation because your phrasing makes me suspect an engineering one. With concepts usually being the first to be culled in a course targeted at engineers, there could be quite a bit of concept debt to pay off before I could really offer something I could honestly call an explanation.
Have you tried the 3Blue1Brown video on the topic[1]? It does not AFAIR offer any answers as to why the Fourier transform should exist or be useful, but it does show very well what it does in the immediate sense.
Like, the other day I learned[0] that if you shine a light through a small opening, the diffraction pattern you get on the other side is basically the Fourier transform of the aperture outline.
(Yes, this also implies that if you take a Fourier transform of an image and make a diffraction grating off the resulting pattern, projecting light through it should paint you the original image.)
--
You can project reality onto any complete basis of functions you like, but this one tends to diagonalize the physics of our universe, which is an overpowered ability inside of our universe.
Because it diagonalizes all good translationally invariant operators, and our universe is fond of translation invariance until you get into general relativity. (This sounds less mysterious once you learn that all good translationally invariant operators are essentially convolutions. Neither of these statements is often taught at the elementary level, probably because of the difficulty and ambiguity in defining “all good” and “are essentially”.)
I first learned this seemingly obvious in hindsight corollary from another comment on HN [1] and it blew my mind. I wish it was included in the usual descriptions of why we choose complex exponential basis for things like the Laplace transform. It's all well and good that they're eigenfunctions of translations, but it still left me wondering why we care about that in the first place.
(If your first exposure is instead from a physics or EE perspective, I suspect the framing would be more obvious, as compared to how it's usually introduced in DiffEq when the choice of basis just seems like a "neat" trick given that it behaves well under differentiation).
You don’t need the convolution statement to see that (as the comment you linked above also demonstrates). A good linear algebra course should have the statement that any set of commuting operators has a common eigenbasis[1]. In particular, if an operator has nondegenerate eigenvalues, then its (essentially unique) eigenbasis is also an eigenbasis for any operator that commutes with it. Take a translation as the former operator and any translation-invariant operator as the latter and you see why all of these just got diagonalized simultaneously.
Above I’ve blatantly ignored all the infinite-dimensional problems that arise when attempting to explain grown-up Fourier transforms, but literally this is actually enough if your Fourier transform is finite-dimensional—the Fourier transform on Z/nZ (aka the discrete Fourier transform on a circle) is most commonly used in applications, but literally everything goes through word for word on an arbitrary finite Abelian group. If your mental powers of abstraction feel like they should be able to acquire some intuition about the real case from the finite case, I highly recommend you read Paul Garrett’s note on the topic[2].
That said, yes, the statement on convolution operators is unreasonably hard to find or stumble upon in the literature. Part of it is that stating it properly is annoying and fussy. Another part is purely terminological: the term “convolution operator” is really rare among books younger than half a century. The usual term is instead “Fourier multiplier” or just “multiplier”, which basically makes sense iff you already know the convolution theorem. Searching for the modern term should give you a plethora of sources. (AFAIU, part of the motivation for this switch is that working with the Fourier transform of your convolution kernel instead of the kernel itself allows one to avoid distributions / generalized functions—and the associated hardcore functional analysis—longer. Consider that, if you want to use kernels, already the literal identity operator forces you to work with the Dirac delta.)
[1] If A and B commute, then any eigenspace ker(A-tE) of A is invariant under B, so decompose your space into a direct sum of eigenspaces of one of your operators and recurse on each. Choose an arbitrary basis when the operators run out.
[2] https://www-users.cse.umn.edu/~garrett/m/repns/notes_2014-15...
> Fourier features -- dimensions in the hidden state that represent numbers via a set of features sparse in the frequency domain
Brilliant and very rigorous!
With no other information, those would be my guesses as to why one would use a GPT2 model.
what's the convention on the meaning of "pre-training" vs "training from scratch" ?
Is this a nomenclature shift?
training from scratch would be initializing a neural network, and training it to add numbers directly.
(Either that, or you linked to the wrong commentary. https://www.lesswrong.com/posts/E7z89FKLsHk5DkmDL/language-m... would be closer, other than being a different paper.)