Maybe it's been done, but I'd like to see an LLM recreate Euclid from questioning without having seen it during training.
Maybe it's been done, but I'd like to see an LLM recreate Euclid from questioning without having seen it during training.
Yes, I believe that, it's part of what I was implying (I believe the LLM weights have some internal representation of math in the same way brains do that allow them to produce proofs)
I think we differ on what "mathematical intuition" is then. I've seen people that do well in undergrad math degrees simply by massively memorising things and learning how to join them up to some level of degrees-of-separation, but seemingly completely fail to understand, for ezample, why even calculus is how it is. Because they are able to regurgitate the results and "produce proofs" this is never questioned.
The Euclid example also shows my bias towards spatial intuition of mathematical concepts (which is deeply unfashionable) but also exposes exactly where at least current LLMs break down; they do the symbol based pattern matching version, but they cannot leap outside of that, at least today.
If you want to catch them, surely you can find proofs they aren't able to produce.
I used to be a game dev, and one of the interview questions someone came up with consisted of working out the surface area of a variant of Menger sponge to some given level of depth. The bifurcation for people that could do this vs those that couldn't was incredible, and did not follow obvious trends for academic achievement. (The same interview also included the gem "How wide is a pointer?" which also catches a frightening number of people).
There is an idea that human intuition, expertise, and critical thinking are largely pattern recognition. When you encounter a situation, your brain gives you a plausible starting point, based on what it has experienced before. You then continue with explicit reasoning, which is slow and inefficient, and try to validate your ideas. The more relevant the patterns you have learned are to the situation, the more likely you reach a useful conclusion.
LLMs are largely the same, except that they cannot learn from experience in normal usage. And except that they experience the world only through symbolic data, while the human brain has access to plenty of sensory data.
Maybe LLMs do not need intuition because they can scale their “cognitive capacity” with hardware and brute force their way through these problem spaces.
My view is that is certainly true of smaller LLMs but becomes less true as they scale up.
To quote the parent bananaflag in a sub-comment:
> I believe the LLM weights have some internal representation of math in the same way brains do that allow them to produce proofs
I think as the sort of spare space adjacent to pure language processing in LLMs grows the probability of the sort of reasoning bananaflag is getting at (or spatial reasoning, or anything else) emerging in that space grows enormously.
One of the questions for AI development over the coming months or years is going to be if deliberately cultivating the architecture of those sub models for specific reasoning types beats any emergent reasoning mechanisms or not.
But to me that is analogous to what human brains do, and a bit different from intuition. I think of intuition as “heuristics”, typically developed through experience, that may link seemingly unrelated concepts via vague, hard-to-define associations, but which let us make mental leaps (or shortcuts) while reasoning. (Maybe analogous to System 1 / 2 thinking.)
On the other hand, LLMs can do both: build “intuition” from patterns in data AND brute force a huge amount of potentially unrelated concepts. This gets fuzzier when we realize that even these “concepts” themselves are gleaned from patterns in data! But my point is we necessarily have to take shortcuts to scale, whereas machines can scale with hardware.
This is of course a layman theory! But it could explain why these models are progressing so fast.
With the alternate view of intuition that many of you are describing it is clear LLMs are somewhat either there or heading there now.
One thing that struck me from Dario's last podcast with Dwarkesh was that he said training LLMs on a diverse set of tasks does not make them better just at those tasks, but they get better at unrelated and other tasks overall. What you described could be a concrete example of how that dynamic works!
Proof is the rigorous outcome.
LLM running probabilistic loop is different kind of process.
Therefore, according to that logic, an entity producing proofs must have intuition.
Edit to add: the parent commenter has now confirmed my interpretation of their statement.
That premise seems unlikely to be correct.
If it's possible for a machine to produce a proof without intuition then clearly a human could also do it too. (And in fact I'd argue I've seen many people like that, simply very good at pattern matching over memorised items).
That doesn’t really read as good faith engagement in the discussion. At best, it reads as being so AI pilled that you can’t even fathom that others might want to have a little side discussion about something other than AI.
What is up with this whole sub thread of obvious hole digging?
Plus, I studied math, I am from that environment. His description matches how math is done by people.
People who are good at pattern matching and memorize are, frankly, shit mathematicians. They are find in fun culture around math, but rarely in actual math. They cant really do it as science.
Well yeah they do, obviously.
LLM cannot reinvent Euclid from scratch, but a larger system including LLM might.