I honestly don't know personally either way. Based on my limited understanding of how LLMs work, I don't see them be making the next great song or next great book and based on that reasoning I'm betting that it probably wont be able to do whatever next "Descartes, Newton, Leibnitz, Gauss, Euler, Ramanujan, Galois" are going to do.
Of course AI as a wider field comes up with something more powerful than LLM that would be different.
Meanwhile, songs are hitting number one on some charts on Spotify that people think are humans and are actually AI. And Spotify has to start labelling them as such. One AI "band" had an entire album of hits.
Also - music is a subjective. Mathematics isn't.
And in this case, an LLM discovered a new way to reason about a conjecture. I don't know how much proof is needed - since that is literally proof that it can be done.
There is quite some questions around that. Music is subjective and obviously different people have different taste, but I wouldn't call any of them to be actual good music / real hits.
>> LLM discovered a new way to reason about a conjecture
I wasn't questioning LLMs ability to prove things. Parent threads were talking about building new kind of maths , or approaching it in a creative/artistic way. Thats' what I was referring to.
I can't speak for maths of hard science as I'm not trained in that, but the creativity aspect in code is definitely lacking when it comes to LLMs. May not matter down the line.
because I have no basis for assuming an LLM is fundamentally capable of doing this.
"Never shall I be beaten by a machine!”
In 1997 he lost to Deep Blue.
The differences between them are many, but brute force doesn't enter into it in either case.
Not a good argument for turning everything over to the Deep Blues. What's Deep Blue done for me lately?
Train an LLM only on texts dated prior to Newton and see if it can create calculus, derrive the equations of motion, etc.
If you ask it about the nature of light and it directs you to do experiments with a prism I'd say we're really getting somewhere.
[1] Obviously Newton counts as one. Leibniz like Newton figured out calculus. Other people did important work in dynamics though no one else's was as impressive as Newton's. But the vast majority of human-level intelligences trained on texts prior to Newton did not create calculus or derive the equations of motion or come close to doing either of those things.
Why are they not coming up with paradigm shift in knowledge expression/discovery like humans did back then?
Are we just not prompting them right?
If we believe today's models are sufficiently capable to have been able to do so, why are we not getting these types of results today compared to the entire world knowledge and especially math?
Are research mathematicians simply not prompting LLMs in the right way?
Incidentally, similar conversations were had about ML writ large vs. classical statistics/methods, and now they've more or less completely died down since it's clear who won (I'm not saying classical methods are useless, but rather that it's obvious the naysayers were wrong). I anticipate the same trajectory here. The main difference is that because of the nature of the domain, everyone has an opinion on LLM's while the ML vs. statistics battle was mostly confined within technical/academic spaces.
What example is there where an LLM has extrapolated? All I've seen is a data set so large and an extra decomposition process making it so interpolation feels like extrapolation if you don't look close enough.
> but a theory of why further advancements can't solve the deficiencies
How about LeCun's?
But if you actually try to take a convex hull of, some encoding of sentences as vectors? It isn’t true. The outputs are not in the convex hull of the training data.
I guess it’s supposed to be a metaphor and not literal, but in that case it’s confusing. Especially seeing as there are contexts in machine learning where literal interpolation vs literal extrapolation, is relevant. So, please, find a better way to say it than saying that “it can only interpolate”?
If it can only interpolate in a literal sense, that means that it only produces good outputs on convex combinations of inputs that appear in the training set. That's what interpolation means. But, if you take the embedding vectors of sentences/prompts, and then take the convex hull of these, it is not typical for new sentences not in the training set to have its embedding vectors be in the convex hull of these.
And, my point is that the inputs it is often fed are not in the convex hull of the inputs in the training data.
When the input space is very high dimensional, this is a common outcome.
I’m not denying that the outputs are causally downstream from the training data. Of course it is.
I’m saying that the inference time inputs aren’t in the convex hull of the training time inputs. This isn’t about saying that the output isn’t because of the training data. Of course it is.
But when you have very high dimensional input space, then even with many inputs in the training data, it is still common for inference time inputs to not be in the convex hull of the train time inputs.
This has nothing to do with the complexities of how the models work after the initial embedding of the tokens as vectors. It’s just about the inputs that appear during training, and the inputs that appear at inference time.
> But an LLM can not infer a concept to which it has no information channel.
Of course! And nothing I said implies otherwise. Really, the point I’m making doesn’t even depend on what the model outputs!
If I took a best fit line from 1 parameter to a 1D output, and then provided that linear model an output that was outside the range of inputs the best fit line was obtained from, that would not be interpolation, it would be extrapolation.
It is similar here, except instead of the input being outside the convex hull due to being further away, it is outside the convex hull due to, like, the shape of the convex hull of training inputs just doesn’t include the point in question.
In the end, creativity has always been a combination of chance and the application of known patterns in new contexts.
If you know anything about the invention of new math (analytic geometry, Calculus, etc.), you'd know how untrue this is. In fact, Calculus was extremely hand-wavy and without rigorous underpinnings until the mid 1800s. Again: more art than science.
If anything, they were fighting an uphill battle against the perception of hand-waving by their contemporaries.
That idea wasn’t formally defined until 134 years later with epsilon-delta by Cauchy. That it was accepted. (I know that there were an earlier proofs)
There’s even arguments that the limit existed before newton and lebnitz with Archimedes' Limits to Value of Pi.
Cauchy’s deep understanding of limits also led to the creation of complex function theory.
These forms of creation are hand-wavy not because they are wrong. They are hand wavy because they leverage a deep level of ‘creative-intuition’ in a subject.
An intuition that a later reader may not have and will want to formalize to deepen their own understanding of the topic often leading to deeper understanding and new maths.
Yes, and it's pretty common knowledge that Calculus was (finally) formalized by Weierstrass in the early 19th century, having spent almost two centuries in mathematical limbo. Calculus was intuitive, solved a great class of problems, but its roots were very much (ironically) vibes-based.
This isn't unique to Newton or Leibniz, Euler did all kinds of "illegal" things (like playing with divergent series, treating differentials as actual quantities, etc.) which worked out and solved problems, but were also not formalized until much later.
Vibe-what? Vibe-bullshit, maybe; cathedrals in Europe and such weren't built by magic. Ditto with sailing and the like. Tons of matematics and geometry there, and tons of damn axioms before even the US existed.
Heck, even the Book of The Games from Alphonse X "The Wise" has both a compendia of game rules and even this https://en.wikipedia.org/wiki/Astronomical_chess where OFC being able on geometry was mandatory at least to design the boards.
On Euclid:
https://en.wikipedia.org/wiki/Euclid%27s_Elements
PD: Geometry has tons of grounds for calculus. Guess why.
Americans and British geeks/nerds are blinded down by Newton unable to realize that there was tons of previous work since the Greek and in Middle Ages, where the British love to depict as brutish people with no culture at all.
And the case is that they weren't dumb at all and without Euclid and Archimede there woudn't be any Calculus.
LLMs are prompted by humans and the right query may make it think/behave in a way to create a novel solution.
Then there's a third factor now with Agentic AI system loops with LLMs. Where it can research, try, experiment in its own loop that's tied to the real world for feedback.
Agentic + LLM + Initial Human Prompter by definition can have it experiment outside of its domain of expertise.
So that's extending the "LLM can't create novel ideas" but I don't think anyone can disagree the three elements above are enough ingredients for an AI to come up with novel ideas.
That's not creative prompt. That's a driving prompt to get it to start its engine.
You could do that nowadays and while it may spend $1,000 to $100,000 worth of tokens. It will create something humans haven't done before as long as you set it up with all its tool calls/permissions.
It won't because even though it looks clever to you, people who /do/ understand math and LLMs understand that LLMs /are/ regurgitating
Why does your LLM need you to tell it to look in the first place? Why isn't just telling us all the answers to unsolved conjectures known and unknown?
Why isn't the LLM just telling us all the answers to all the problems we are facing?
Why isn't the LLM telling us, step by step with zero error, how to build the machine that can answer the ultimate question?
> Timothy Gowers @wtgowers
> @wtgowers
> If you are a mathematician, then you may want to make sure you are sitting down before reading further.
If your refutation requires someone to have an account, login, and read something - it's meaningless
it's readable to most, it's annoying having to swamp through ex-Twitter .. but there are work around's.
But, I remain sceptical
https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29a...
it includes the longer remarks by Gowers & others.
We just haven't let AI run wild yet. But its coming.
AGI has been "just over the horizon" for literal decades now - there have been a number of breakthroughs and AI Winters in the past, and there's no real reason to believe that we've suddenly found the magic potion, when clearly we haven't.
AI right now cannot even manage simple /logic/
Who decides at which the last point it’s OK to provide text to the model in order to be able to describe it as creative? (non-rhetorical)