Can AI do maths yet? Thoughts from a mathematician
xenaproject.wordpress.com
xenaproject.wordpress.com
I'm a research mathematician. In the 1980's I'd ask everyone I knew a question, and flip through the hard bound library volumes of Mathematical Reviews, hoping to recognize something. If I was lucky, I'd get a hit in three weeks.
Internet search has shortened this turnaround. One instead needs to guess what someone else might call an idea. "Broken circuits?" Score! Still, time consuming.
I went all in on ChatGPT after hearing that Terry Tao had learned the Lean 4 proof assistant in a matter of weeks, relying heavily on AI advice. It's clumsy, but a very fast way to get suggestions.
Now, one can hold involved conversations with ChatGPT or Claude, exploring mathematical ideas. AI is often wrong, never knows when it's wrong, but people are like this too. Read how the insurance incidents for self-driving taxis are well below the human incident rates? Talking to fellow mathematicians can be frustrating, and so is talking with AI, but AI conversations go faster and can take place in the middle of the night.
I don't want AI to prove theorems for me, those theorems will be as boring as most of the dreck published by humans. I want AI to inspire bursts of creativity in humans.
Why do I need to hire an artist for my movie/video game/advertisement when AI can replicate all the creativity I need.
https://direct.mit.edu/rest/article-abstract/102/3/583/96779...
It would be interesting if in the future mathematicians are just as fluent in some (possibly AI-powered) proof verifying tool, as they are with LaTeX today.
Here it means do math research or better, find new math.
More problematic with that statement is that a timeline isn’t specified. 1 year? Probably not. 10 years? Probably. 20 years? Very likely. 100 years? None of us here will be alive to be proven wrong but I’ll venture that that’s a certainty.
I will agree that it's likely none of us here will be alive to be proven wrong, but that's in the 1 to 10 year range.
I don’t see how my position is so exceptionally strong. I’m saying indeed there’s a 55-70% probability that this happens in the 1-10 year time frame. At 1-20 it goes up to 70-90%. It’s still important to leave room for doubt that we might miss something or be unable to build something for a long time. Trying to state otherwise seems like an even stronger position to take to me.
We are still early in AI.
When talking with various models of ChatGPT about research math, my biggest gripe is that it's either confidently right (10% of my work) or confidently wrong (90%). A human researcher would be right 15% of the time, unsure 50% of the time, and give helpful ideas that are right/helpful (25%) or wrong/a red herring (10%). And only 5% of the time would a good researcher be confidently wrong in a way that ChatGPT is often.
In other words, ChatGPT completely lacks the meta-layer of "having a feeling/knowing how confident it is", which is so useful in research.
Just assume that the LLM is wrong and test their assumptions - math is one of the few disciplines where you can do that easily
To clarify, I was doing research on applied math. My field is not analysis, but I needed to prove some bounds on certain messed up expressions (involving special functions, etc), and analyze an ODE that's not analytically solvable. I used the COT model a fair bit.
I would ask ChatGPT for hints/ideas/direction in proving various bounds, asking it for theorems or similar results in literature. This is exactly the kind of thing where a researcher would go "yeah this looks like X" or "I think I saw something like this in (book/article name)", or just know a method; or alternatively say they have no clue. ChatGPT most often will confidently give me a "solution", being right 10% of the time (when there's a pretty standard way to do it that I didn't see/know).
On the whole it was quite useful.
I don't think AI will think conventionally. It isn't thinking to begin with. It is weighing options. Those options permutate and that is why every response is different.
PS: I'm starting to see a lot of plausible deniability in some comments about LLMs capabilites. When LLMs do great => "cool, we are scaling AI". when LLMs do something wrong => "user problem", "skill issues", "don't judge a fish for its ability to fly".
In any case my best experiences with LLMs for pure math research have been for exploring the problem space and ideation -- queries along the line of "Here's a problem I'm working on ... . Do any other fields have a version of this problem, but framed differently?" or "Give me some totally left field methods, even if they are from different fields or unlikely to work. Assume I've exhausted all the 'obvious' approaches from field X"
Of course they are, I hoped it was clear I was just sharing my experience trying to use it for research!
I did in general word it as I would a question to a researcher, which includes an uncertainty in it being true. E.g. this is from a recent prompt: "is this true in general, if not, what are the conditions for this to be true?"
I think this is a sensible tuning in that it's probably what most people who log on to chatgpt want. Most questions people ask of it will have simple enough answers that require knowledge but not all that much reasoning.
But I see no reason why it couldn't be tuned to be more open ended, less eager to give the correct benchmark/exam answer right away. Indeed in the "internal narrative" of recent models, I see them ask themselves things I wish they asked me!
I made a slight change to generalise your statement, I think you have summarised the actual marketing opportunity.
Improved tooling and techniques have given humans the free time and resources needed for arts, culture, philosophy, sports, and spending time to enjoy life! Fancy telecom technologies have allowed me to work from home and i love it :)
I find that a lot of AI+Math work is focused on the end game where you have a clear problem to solve, rather than the early exploratory work where most of the time is spent. The challenge is in making the right connections and analogies, discovering hidden useful results, asking the right questions, translating between fields.
I'm getting ready to launch [Sugaku](https://sugaku.net), where I'm trying to build tools for the above, based on processing the published math literature and training models on it. The kind of search of MR that you mentioned doing is exactly what a computer should do instead. I can create an account for you and would love some feedback.
Also Glazer seemed to regret calling T1 "IMO/undergraduate", and not only because of the disparity between IMO and typical undergraduate. He said that "We bump problems down a tier if we feel the difficulty comes too heavily from applying a major result, even in an advanced field, as a black box, since that makes a problem vulnerable to naive attacks from models"
Also, all of the problems shows to Tao were T3
that's if you unconditionally believe in result without any proofreading, confirmation, reproducability and even barely any details (we are given only one slide).
The "reality" of keeping this stuff secret 'cause someone would train on it is itself bizarre and certainly shouldn't be above questioning.
https://www.reddit.com/r/OpenAI/comments/1hiq4yv/comment/m30...
That is the "reality" - that because companies can train their models on the whole Internet, companies will train their (base) models on the entire Internet.
And in this situation, "having heard the problem" actually serves as a barrier to understanding of these harder problems since any variation of known problem will receive a standard "half-assed guestimate".
And these companies "can't not" use these base models since they're resigned to the "bitter lesson" (better the "bitter lesson viewpoint" imo) that they need large scale heuristics for the start of their process and only then can they start symbolic/reasoning manipulations.
But hold-up! Why couldn't an organization freeze their training set and their problems and release both to the public? That would give us an idea where the research stands. Ah, the answer comes out, 'cause they don't own the training set and the result they want to train is a commercial product that needs every drop of data to be the best. As Yan LeCun has said, this isn't research, this is product development.
Don't kid yourself. There are 10's of billions of dollars going into AI. Some of the humans involved would happily cheat on comparative tests to boost investment.
Having a higher valuation could help with attracting better talent or more funding to invest in GPUs and actual model improvements but I don't think that outweighs the risks unless you're a tiny startup with nothing to show (but then you wouldn't have the money to bribe anyone).
It depends a lot on individuals making up the companies command chain and their values.
CEOs and VCs will happily lie because they are convinced they are smarter than everyone else and will solve the problem before they get caught.
O1 is a lot better at spotting its errors than 4o but it too still makes a lot of really stupid mistakes. It seems to be quite far from producing results itself consistently without at least a somewhat clueful human doing hand-holding.
It never tells you the wrong thing, at the very least.
Hopefully they have incorporated more modern LLM since then, but it hasn’t been that long.
You can unlock a full derivation of the solution, for cases where you say "Solve" or "Simplify", but what I (and I suspect GP) might want, is to know why a few of the key steps might work.
It's a fantastic tool that helped get me through my (engineering) grad work, but ultimately the breakthrough inequalities that helped me write some of my best stuff were out of a book I bought in desperation that basically cataloged linear algebra known inequalities and simplifications.
When I try that kind of thing with the best LLM I can use (as of a few months ago, albeit), the results can get incorrect pretty quickly.
https://www.amazon.com/gp/product/048667102X/ref=ppx_yo_dt_b...
I make no claim about its usefulness for anyone else!
It's been some time since I've used the step-by-step explainer, and it was for calculus or intro physics problems at best, but IIRC the pro subscription will at least mention the method used to solve each step and link to reference materials (e.g., a clickable tag labeled "integration by parts"). Doesn't exactly explain why but does provide useful keywords in a sequence that can be used to derive the why.
Of course there would be the risk of adversaries giving bogus feedback, but my gut says it's relatively straightforward to filter out most of this muck.
I was figuring out some mode decomposition methods such as ESPRIT and Prony and how to potentially extend/customize them. Wolfram Alpha doesn't seem to have a clue about such.
WolframOne/Mathematica is better, but that requires the user (or ChatGPT!)to write complicated code, not natural language queries.
For example I asked Wolfram Alpha "How heavy a rocket has to be to launch 5 tons to LEO with a specific impulse of 400s", which is a straightforward application of the Tsiolkovsky rocket equation. Wolfram Alpha gave me some nonsense about particle physics (result: 95 MeV/c^2), GPT-4o did it right (result: 53.45 tons).
Wolfram alpha knows about the Tsiolkovsky rocket equation, it knows about LEO (low earth orbit), but I found no way to get a delta-v out of it, again, more nonsense. It tells me about Delta airlines, mentions satellites that it knows are not in LEO. The "natural language" part is a joke. It is more like an advanced calculator, and for that, it is great.
It dates back to Steve Jobs blaming an iPhone 4 user for "holding it wrong" rather than acknowledging a flawed antenna design that was causing dropped calls. The closest Apple ever came to admitting that it was their problem was when they subsequently ran an employment ad to hire a new antenna engineering lead. Maybe it's time for Wolfram to hire a new language-model lead.
The problem has always been that you only get good answers if you happen to stumble on a specific question that it can handle. Combining Alpha with an LLM could actually be pretty awesome, but I'm sure it's easier said than done.
The GP's rocket equation question is exactly the sort of use case for which Alpha has been touted for years.
Has anyone had any luck with this? It seems like the only thing that it just can't do.
Ask it to produce graphs with python and matplotlib. That will work.
When you visit chatgpt on the free account it automatically gives you the best model and then disables it after some amount of work and says to come back later or upgrade.
Giving AI the ability to execute code is the safety peoples nightmare though, wonder if we'll hear anything from them as this is surely coming
It often gets the actual math wrong, but it is good enough at connecting the dots between my layman's intuition and the "right answer" that I can get myself over humps that I'd previously have been hopelessly stuck on.
It does make those mistakes you're talking about very frequently, but once I'm told that the thing I'm trying to do is achievable with the Gram-Schmidt process, I can go self-educate on that further.
The big thing I've had to watch out for is that it'll usually agree that my approach is a good or valid one, even when it turns out not to be. I've learned to ask my questions in the shape of "how do I", rather than "what if I..." or "is it a good idea to...", because most of the time it'll twist itself into shapes to affirm the direction I'm taking rather than challenging and refining it.
[ (Re)imagining mathematics in a world of reasoning machines by Akshay Venkatesh]
https://www.youtube.com/watch?v=vYCT7cw0ycw [54min]
Abstract: In the coming decades, developments in automated reasoning will likely transform the way that research mathematics is conceptualized and carried out. I will discuss some ways we might think about this. The talk will not be about current or potential abilities of computers to do mathematics—rather I will look at topics such as the history of automation and mathematics, and related philosophical questions.
See discussion at https://news.ycombinator.com/item?id=42465907
"We might put the axioms into a reasoning apparatus like the logical machinery of Stanley Jevons, and see all geometry come out of it. That process of reasoning are replaced by symbols and formulas... may seem artificial and puerile; and it is needless to point out how disastrous it would be in teaching and how hurtful to the mental development; how deadening it would be for investigators, whose originality it would nip in the bud. But as used by Professor Hilbert, it explains and justifies itself if one remembers the end pursued." Poincare on the value of reasoning machines, but the analogy to mathematics once we have theorem-proving AI is clear (that the tools and the lie direct outputs are not the ends. Human understanding is).
"Even if such a machine produced largely incomprehensible proofs, I would imagine that we would place much less value on proofs as a goal of math. I don't think humans will stop doing mathematics... I'm not saying there will be jobs for them, but I don't think we'll stop doing math."
"Mathematics is the study of reproducible mental objects." This definition is human ("mental") and social (it implies reproducing among individuals). "Maybe in this world, mathematics would involve a broader range of inquiry... We need to renegotiate the basic goals and values of the discipline." And he gives some examples of deep questions we may tackle beyond just proving theorems.
But I'm wondering what other people think of this analogy.
I used to be a bench scientist (molecular genetics).
There were world class researchers who were more creative than I was. I even had a Nobel Laureate once tell me that my research was simply "dotting 'i's and crossing 't's".
Nevertheless, I still moved the field forward in my own small ways. I still did respectable work.
So, will these LLMs make us completely obsolete? Or will there still be room for those of us who can dot the "i"?--if only for the fact that LLMs don't have infinite time/resources to solve "everything."
I don't know. Maybe I'm whistling past the graveyard.
Nobel laureate and winner are the same thing.
> Linus Pauling was talking absolute garbage, harmful and evil, after winning the Nobel.
Can you be more specific, what garbage? And which Nobel prize do you mean – Pauling got two, one for chemistry and one for peace.
https://scarc.library.oregonstate.edu/coll/pauling/blood/nar...
The goal wasn't to mark people for ostracism but to make it easier for people carrying these genes to find mates that won't result in suffering for their offspring.
Not a fun thing to discuss, but apparently a significant issue, which I guess should be unsurprising given some of the laws allowing underage marriage if the family signs off.
Mentioning only to draw attention to the fact that theoretical policy is often undeniable in a vacuum, but runs aground when faced with real world conditions.
I was referring to Linus's harmful and evil promotion of Vitamin C as the cure for everything and cancer. I don't think Linus was attaching that garbage to any particular Nobel prize. But people did say to their doctors: "Are you a Nobel winner, doctor?". Don't think they cared about particular prize either.
Which is "harmful and evil" thanks to your afterknowledge. He had based his books on the research that failed to replicate. But given low toxicity of vitamin C it's not that "evil" to recommend treatment even if probabilistic estimation of positive effects is not that high.
Sloppy, but not exceptionally bad. At least it was instrumental in teaching me to not expect marvels coming from dietary research.
Digression aside, my point is that I don’t think we know exactly what makes or defines “the golden hands”. And if that is the case, can we optimize for it?
Another point is that scalable fine tuning only works for verifiable stuff. Think a priori knowledge. To me that seems to be at the opposite end of the spectrum from “mess with it and see what happens”.
Very funny. My friends and I never used the phrase "golden hands" but we used to say something similar: "so-and-so has 'great hands'".
But it meant the same thing.
I, myself, did not have great hands, but my comment was more about the intellectual process of conducting research.
I guess my point was that:
* I've already dealt with more talented researchers, but I still contributed meaningfully.
* Hopefully, the "AI" will simply add another layer of talent, but the rest of us lesser mortals will still be able to contribute.
But I don't know if I'm correct.
The best part of math (again, just for me) is that it was a journey that was done by hand with only the human intellect that computers didn't understand. The beauty of the subject was precisely that it was a journey of human intellect.
As I said elsewhere, my friends used to ask me why something was true and it was fun to explain it to them, or ask them and have them explain it to me. Now most will just use some AI.
Soulless, in my opinion. Pure mathematics should be about the art of the thing, not producing results on an assembly line like it will be with AI. Of course, the best mathematicians are going into this because it helps their current careers, not because it helps the future of the subject. Math done with AI will be a lot like Olympic running done with performance-enhancing drugs.
Yes, we will get a few more results, faster. But the results will be entirely boring.
For myself, chasing lemmas was always boring — and there’s little interest in doing the busywork of fleshing out a theory. For me, LLMs are a great way to do the fun parts (conceptual architecture) without the boring parts.
And I expect we’ll such much the same change as with physics: computers increase the complexity of the objects we study, which tend to be rather simple when done by hand — eg, people don’t investigate patterns in the diagrams of group(oids) because drawing million element diagrams isn’t tractable by hand. And you only notice the patterns in them when you see examples of the diagrams at scale.
The the trick is teaching the thing how high powered of theorems to use or how to factor out details or not depending on the user's level of understanding. We'll have to find a pedagogical balance (e.g. you don't give `linarith` to someone practicing basic proofs), but I'm sure it will be a great tool to aid human understanding.
A tool to help translate natural language to formal propositions/types also sounds great, and could help more people to use more formal methods, which could make for more robust software.
As Erdos(I think?) said, great math is not about the answers, it's about the questions. Or maybe it was someone else, and maybe "great mathematicians" rather than "great math". But, gist is the same.
"What happens when you invent a thing that makes a function continuous (aka limit point)"? "What happens when you split the area under a curve into infinitesimal pieces and sum them up"? "What happens when you take the middle third out of an interval recursively"? "Can we define a set of axioms that underlie all mathematics"? "Is the graph of how many repetitions it takes for a complex number to diverge interesting"? I have a hard time imagining computers would ever have a strong enough understanding of the human experience with mathematics to even begin pondering such questions unprompted, let alone answer them and grok the implications.
Ultimately the truths of mathematics, the answers, soon to be proved primarily by computers, already exist. Proving a truth does not create the truth; the truth exists independent of whether it has been proved or not. So fundamentally math is closer to archeology than it may appear. As such, AI is just a tool to help us dig with greater efficiency. But it should not be considered or feared as a replacement for mathematicians. AI can never take away the enlightenment of discovering something new, even if it does all the hard work itself.
The key is that the good questions however come from hard-won experience, not lazily questioning an AI.
https://www.wired.com/story/defeated-chess-champ-garry-kaspa...
If you believe the purpose of pure math is to shed light on patterns in nature, pave the way for the sciences, etc., this is fantastic news.
I could see how AI could assist me with learning pure math but the idea AI is going to do pure math for me is just absurd.
Not only would I not know how to start, more importantly I have no interest in pure math. There will still be a huge time investment to get up to speed with doing anything with AI and pure math.
You have to know what questions to ask. People with domain knowledge seem to really be selling themselves short. I am not going to randomly stumble on a pure math problem prompt when I have no idea what I am doing.
Do people with PhD in math really ask AI to explain math concepts to them?
> Now most will just use some AI.
I'm genuninely wondering whether it's true. The "now" and "most" part.
Whatever they write only happens to contain some truth by virtue of the model and the training data. An algorithm doesn’t know what truth is or why we value it. It’s a bullshitter of the highest calibre.
Then comes the question: will they write proofs that we will consider beautiful and elegant, that we will remember and pass down?
Or will they generate what they’ve been asked to and nothing less? That would be utterly boring to read.
LLMs will be tools for some math needs and even if we ever get quantum computers will be limited in what they can do.
LLMs, without pattern matching, can only do up to about integer division, and while they can calculate parity, they can't use it in their calculations.
There are several groups sitting on what are known limitations of LLMs, waiting to take advantage of those who don't understand the fundamental limitations, simplicity bias etc...
The hype will meet reality soon and we will figure out where they work and where they are problematic over the next few years.
But even the most celebrated achievements like proof finding with Lean, heavily depends on smart people producing hints that machines can use.
Basically lots of the fundamental hints of the limits of computation still hold.
Model logic may be an accessable way to approach the limits of statistical inference if you want to know one path yourself.
A lot of what is in this article relates to some the known fundamental limitations.
Remember that for all the amazing progress, one of the core founders of the perceptron, Pitts drank him self to death in the 50s because it was shown that they were insufficient to accurately model biological neurons.
Optimism is high, but reality will hit soon.
So think of it as new tools that will be available to your child, not a replacement.
It is an effect of the complex to unpack descriptive complexity class DLOGTIME-uniform TC0, which has AND, OR and MAJORITY gates.
http://arxiv.org/abs/2409.13629
The point being that the ability to use parity gates is different than being able to calculate it, which is where the union of the typically ram machine DLOGTIME with the circuit complexity of uniform TC0 comes into play.
PARITY, MAJ, AND, and OR are all symmetric, and are in TCO, but PARITY is not in DLOGTIME-uniform TC0, which is first-order logic with Majority quantifiers.
Another path, if you think about symantic properties and Rice's theorem, this may make sense especially as PAC learning even depth 2 nets is equivalent to the approximate SVP.
PAC-learning even depth-2 threshold circuits is NP-hard.
https://www.cs.utexas.edu/~klivans/crypto-hs.pdf
For me thinking about how ZFC was structured so we can keep the niceties of the law of the excluded middle, and how statistics pretty much depends on it for the central limit and law of large numbers, IID etc...
But that path runs the risk of reliving the Brouwer–Hilbert controversy.
Thank you for the question.
I guess what I'm saying is:
Will LLMs (or whatever comes after them) be _so_ good and _so_ pervasive that we will simply be able to say, "Hey ChatGPT-9000, I'd like to see if the xyz conjecture is correct." And then ChatGPT-9000 just does the work without us contributing beyond asking a question.
Or will the technology be limited/bound in some way such that we will still be able to use ChatGPT-9000 as a tool of our own intellectual augmentation and/or we could still contribute to research even without it.
Hopefully, my comment clarifies my original post.
Also, writing this stuff has helped me think about it more. I don't have any grand insight, but the more I write, the more I lean toward the outcome that these machines will allow us to augment our research.
Every LLM release moves half of the remaining way to the minimum viable goal of replacing a third class undergrad. If your business or research initiative is fine with that level of competence then you will find utility.
The problem is that I don't know anyone who would find that useful. Nor does it fit within any existing working methodology we have. And on top of that the verification of any output can take considerably longer than just doing it yourself in the first place, particularly where it goes off the rails, which it does all the time. I mean it was 3 months ago I was arguing with a model over it not understanding place-value systems properly, something we teach 7 year olds here?
But the abstract problem is at a higher level. If it doesn't become a general utility for people outside of mathematics, which is very very evident at the moment by the poor overall adoption and very public criticism of the poor result quality, then the funding will dry up. Models cost lots of money to train and if you don't have customers it's not happening and no one is going to lend you the money any more. And then it's moot.
But the main question is still: assuming you replace an undergrad with a model, who checks the work? If you have a good process around that already, and find utility as an augmented system, then get you’ll get value - but I still think it’s better for the undergrad to still have the job and be at the wheel, and does things faster and better when leveraging a powerful tool.
Well to be fair no one checks what the graduates do properly, even if we hired KPMG in. That is until we get sued. But at least we have someone to blame then. What we don't want is something for the graduate to blame. The buck stops at someone corporeal because that's what the customers want and the regulators require.
That's the reality and it's not quite as shiny and happy as the tech industry loves to promote itself.
My main point, probably cleared up with a simple point: no one gives a shit about this either way.
That craving for an understanding an elegant proof is nowhere to be found with verifying an LLM’s proof.
Like sure, you could put together a car by first building an airplane, disassembling all of it minus the two front seats, and having zero elegance and still get a car at the end. But if you do all that and don’t provide novelty in results or useful techniques, there’s no business.
Hell, I can’t even get a model to calculate compound interest for me (save for the technicality of prompt engineering a python function to do it). What do I expect?
At this stage, the point they're making isn't 'OMG AGI!!!' but rather something like 'having an enthusiastic, often wrong undergrad assistant who's available 24/7 can be useful, if you use it carefully.'
This time could be different, of course. But I'll need a lot more evidence before I start telling people to base their major life decisions on projected technological change.
That's before we even consider that only a very slim minority of the people who study math (or physics or statistics or biology or literature or...) go on to work in the field of math (or physics or statistics or biology or literature or...). AI could completely take over math research and still have next to impact on the value of the skills one acquires from studying math.
Or if you want to be more fatalistic about it: if AI is going to put everyone out of work then it doesn't really matter what you do now to prepare for it. Might as well follow your interests in the meantime.
We're all usually (but not always) better off, with more productivity, eventually, but in the meantime, jobs do disappear. Robotics did not fully displace machinists and factory workers, but single-skilled people in Detroit did not do well. The loom, the steam engine... all of them displaced often highly-trained often low-skilled artisans.
We've had changes before, the most recent one being the rise of computers and then the internet, and before that, manufacturing automation. In all cases, some people were better prepared for change, and some less so.
The general consensus is that diverse skills and foundational skills (e.g. math, communication) best prepare people for transitions, relative to specialized skills (e.g. one technology). In addition, many careers are likely to be less impacted, such as plumbing.
Most likely AI will be good at some things and not others, and mathematicians will just move to whatever AI isn't good at.
Alternatively, if AI is able to do all math at a level above PhDs, then its going to be a brave new world and basically the singularity. Everything will change so much that speculating about it will probably be useless.
(。•́︿•̀。)
Don't cry for me, (Argentina).
When "God" tells you that you're merely dotting "i"s and crossing "t"s ... he's kind of correct.
But I was arrogant enough to believe that I was still a good researcher. I just wasn't "God."
Evaluating the output of llms will also require mathematical skills.
But I'd go further, if your son enjoys mathematics and has some ability in the area, it's wonderful for your inner life. Anyone who becomes sufficiently interested in anything will rediscover mathematics lurking at the bottom.
There's never been the equivalent of the 'bench scientist' in mathematics and there aren't many direct careers in mathematics, or pure mathematics at least - so very few people ultimately become researchers. Instead, I think you take your way of thinking and apply it to whatever else you do (and it certainly doesn't do any harm to understand various mathematical concepts incredibly well).
But specifically to your worry about humans just dotting i's and crossing t's, he predicts that exactly the opposite will happen. At the end he emphasizes that the ultimate goal of mathematics is more about human understanding than proving theorems.
Historically, the claim that neural nets were actual models of the human brain and human thinking was always epistemically dubious. It still is. Even as the practical problems of producing better and better algorithms, architectures, and output have been solved, there is no reason to believe a connection between the mechanical model and what happens in organisms has been established. The most important point, in my view, is that all of the representation and interpretation still has to happen outside the computational units. Without human interpreters, none of the AI outputs have any meaning. Unless you believe in determinism and an overseeing god, the story for human beings is much different. AI will not be capable of reason until, like humans, it can develop socio-rational collectivities of meaning that are independent of the human being.
Researchers seemed to have a decent grasp on this in the 90s, but today, everyone seems all too ready to make the same ridiculous leaps as the original creators of neural nets. They did not show, as they claimed, that thinking is reducible to computation. All they showed was that a neural net can realize a boolean function—which is not even logic, since, again, the entire semantic interpretive side of the logic is ignored.
The universal approximation theorem. And that's basically it. The rest is empirical.
No matter which physical processes happen inside the human brain, a sufficiently large neural network can approximate them. Barring unknowns like super-Turing computational processes in the brain.
If it does, well, it will take more time to incorporate those physical mechanisms into computers to get them on par with the brain.
I leave the possibility that it's "magic"[1] aside. It's just impossible to predict, because it will violate everything we know about our physical world.
[1] One example of "magic": we live in a simulation and the brain is not fully simulated by the physics engine, but creators of the simulation for some reason gave it access to computational resources that are impossible to harness using the standard physics of the simulated world. Another example: interactionistic soul.
Sorry, but there's not much evidence that can support human exceptionalism.
Although "non-deterministic" and "stochastic" are often used interchangeably, they are not equivalent. Probability is applied analysis whose objects are distributions. Analysis is a form of deductive, i.e. mechanical, reasoning. Therefore, it's more accurate (philosophically) to identify mathematical probability with determinism. Probability is a model for our experience. That doesn't mean our experience is truly probabilistic.
Humans aren't exceptional. Math modeling and reasoning are human activities.
And physicists regard those as unphysical: the theory breaks down, we need better one.
Switching into a less facetious mode...
Do you understand that in context of this dialogue it's not enough to show some examples of discontinuous or otherwise unrepresentable by NNs functions? You need at least to give a hint why such functions cannot be avoided while approximating functionality of the human brain.
Many things are possible, but I'm not going to keep my mind open to a possibility of a teal Russell's teapot before I get a hint at its existence, so to speak.
The brain is a physical system, so whatever it does (including philosophy) can be replicated by modelling (a (vastly) simplified version of) underlying physics.
Anyway, I am not especially interested in discussing possible impossibility of an LLM-based AGI. It might be resolved empirically soon enough.
Or perhaps, determinism and mechanistic materialism - which in STEM-adjacent circles has a relatively prevalent adherence.
Worldviews which strip a human being of agency in the sense you invoke crop up quite a lot today in such spaces. If you start of adopting a view like this, you have a deflationary sword which can cut down most any notion that's not mechanistic in terms of mechanistic parts. "Meaning? Well that's just an emergent phenomenon of the influence of such and such causal factors in the unrolling of a deterministic physical system."
Similar for reasoning, etc.
Now obviously large swathes of people don't really subscribe to this - but it is prevalent and ties in well with utopian progress stories. If something is amenable to mechanistic dissection, possibly it's amenable to mechanistic control. And that's what our education is really good at teaching us. So such stories end up having intoxicating "hype" effects and drive fundraising, and so we get where we are.
For one, I wish people were just excited about making computers do things they couldn't do before, without needing to dress it up as something more than it is. "This model can prove a set of theorems in this format with such and such limits and efficiency"
The irony is: why would someone want control if they don't have true choice? Unfortunately, such a question rarely pierces the intoxicated mind when this mind is preoccupied with pass the class, get an A, get a job, buy a house, raise funds, sell the product, win clients, gain status, eat right, exercise, check insta, watch the game, binge the show, post on Reddit, etc.
Is this controversial in some way? The problem is that to simulate a universe you need a bigger universe -- which doesn't exist (or is certainly out of reach due to information theoretical limits)
> ---like Leibniz's Ratiocinator. The intoxication may stem from the potential for predictability and control.
I really don't understand the 'control' angle here. It seems pretty obvious that even in a purely mechanistic view of the universe, information theory forbids using the universe to simulate itself. Limited simulations, sure... but that leaves lots of gaps wherein you lose determinism (and control, whatever that means).
It’s not “controversial”, it’s just not a given that the universe is to be thought a deterministic machine. Not to everyone, at least.
My comments are not about simulating the universe on a real machine. They're about the validity and value of math/computational modeling in a universe where determinism is scientifically indeterminable.
What would you say if we can predict the outcome of an experiment with 51% probability. Is that enough to establish what you call "control"? What if we can repeat the experiment as many times as we like?
(I must admit, I still don't really understand what "control" means to you, but let's get the preliminaries out of the way first.)
I don’t think it does. Taking computers as an analogy… if you have a computer with 1GB memory, then you can’t simulate a computer with more than 1GB memory inside of it.
Consider also that even as digital technology and the ratiomathimatical understanding of the world has advanced it is still rife with dynamics and problems that require a humanistic approach. In particular, a mathematical conception cannot resolve teleological problems which require the establishment of consensus and the actual determination of what we, as a species, want the world to look like. Climate change and general economic imbalance are already evidence of the kind of disasters that mount when you limit yourself to a reductionistic, overly mathematical and technological understanding of life and existence. Being is not a solely technical problem.
19th/20th was a golden era of philosophy with a coherent and rigorous mathematical lens to apply with other lenses. Russel, Turing, Godel, etc. However this just doesn't exist anymore
The relevance of mathematics to the cognitive problem must be decided outside of mathematics. As another poster said, even if you buy the theorems, it is still an empirical question as to whether or not they really model what they claim to model, and whether or not that model is of a fidelity that we find acceptable for a definition of general intelligence. Often, people reach claims of adequacy today not by producing really fantastic models but instead by lowering the bar enormously. They claim that these models approximate humans by severely reducing the idea of what it means to be an intelligent human to the specific talents their tech happens to excel at (e.g. apparently being a language parrot is all that intelligence is, ignoring all the very nuanced views and definitions of intelligence we have come up with over the course of history. A machine that is not embodied ina skeletal structure and cannot even experience, let alone solve, the vast number of physical, anatomical problems we contend with on a daily basis is, in my view, still very far from anything I would call general intelligence).
It's worth considering why "everyone seems all too ready to make ... leaps ..." "Neural", "intelligence", "learning", and others are metaphors that have performed very well as marketing slogans. Behind the marketing slogans are deep-pocketed, platformed corporate and government (i.e. socio-rational collective) interests. Educational institutions (another socio-rational collective) and their leaders have on the whole postured as trainers and preparers for the "real world" (i.e. a job), which means they accept, support, and promote the corporate narratives about techno-utopia. Which institutions are left to check the narratives? Who has time to ask questions given the need to learn all the technobabble (by paying hundreds of thousands for 120 university credits) to become a competitive job candidate?
I've found there are many voices speaking against the hype---indeed, even (rightly) questioning the epistemic underpinnings of AI. But they're ignored and out-shouted by tech marketing, fundraising politicians, and engagement-driven media.
1. I asked it to show me the derivation of a formula for the efficiency of Stop-and-Wait ARQ and it seemed to do it, but a day later, I realised that in one of the steps, it just made a term vanish to get to the next step. Obviously, I should have verified more carefully, but when I asked it to spot the mistake in that step, it did the same thing twice more with bs explanations of how the term is absorbed.
2. I asked it to provide me syllogisms that I could practice proving. An overwhelming number of the syllogisms it gave me were inconsistent and did not hold. This surprised me more because syllogisms are about the most structured arguments you can find, having been formalized centuries ago and discussed extensively since then. In this case, asking it to walk step-by-step actually fixed the issue.
Both of these were done on the free plan of ChatGPT, but I can remember if it was 4o or 4.
Since chatgpt-4o, there has been o1-preview, and o1 (full) is out. They just announced o3 got a 25% on frontiermath which is what this article is a reaction to. So, any tests on 4o are at least TWO (or three) AI releases with new capabilities.
That means the questions went over the fence to OpenAI.
I'm quite certain they are aware of that, and it would be pretty foolish not to take advantage of at least knowing what the questions are.
Needless to say, it doesn't bring us any closer to AGI.
The only solution I see here is people crafting their own, private benchmarks that the big players don't care about enough to train on. That, at least, gives you a clearer view of the field.
I'm being completely serious. You are correct, despite the downvotes, that this could not be pushing us towards AGI because if the dataset is leaked you can't claim the G-- generalizability.
The point of the benchmark is to lead is to believe that this is a substantial breakthrough. But a reasonable person would be forced to conclude that the results are misleading to due to optimizing around the training data.
In the same way ChatGPT scores 25% on this and the question is "How close were those 25% to questions in the training set". Or to put it another way we want to answer the question "Is ChatGPT getting better at applying it's reasoning to out-of-set problems or is it pulling more data into it's training set". Or "Is the test leaking into the training".
Maybe the whole question is academic and it doesn't matter, we solve the entire problem by pulling all human knowledge into the training set and that's a massive benefit. But maybe it implies a limit to how far it can push human knowledge forward.
In the past, there was quite a bit of low hanging fruit such that you could have polymaths able to contribute to a wide variety of fields, such as Newton.
But in the past 100 years or so, the problem is there is so much known, it is impossible for any single person to have deep knowledge of everything. e.g. its rare to find a really good mathematician who also has a deep knowledge (beyond intro courses) about say, chemistry.
Would a sufficiently powerful AI / ML model be able to come up with this synthesis across fields?
The challenge is how to figure out if a model is genuinely reasoning
Knowledge creation comes from collecting data from the real world, and cleaning it up somehow, and brainstorming creative models to explain it.
NN/LLM's version of model building is frustrating because it is quite good, but not highly "explainable". Human models have higher explainability, while machine models have high predictive value on test examples due to an impenetrable mountain of algebra.
This is factually wrong. The most interesting problems motivating the quantum computing research are hard to solve, but easy to verify on classical computers. The factorization problem is the most classical example.
The problem is that existing quantum computers are not powerful enough to solve the interesting problems, so researchers have to invent semi-artificial problems to demonstrate "quantum advantage" to keep the funding flowing.
There is a plethora of opportunities for LLMs to show their worth. For example, finding interesting links between different areas of research or being a proof assistant in a math/programming formal verification system. There is a lot of ongoing work in this area, but at the moment signal-to-noise ratio of such tools is too low for them to be practical.
You parent did not talk about quantum computers. I guess he rather had predictions of novel quantum-field theories or theories of quantum gravity in the back of his mind.
> Having said that, the biggest caveat to the “10^25 years” result is one to which I fear Google drew insufficient attention. Namely, for the exact same reason why (as far as anyone knows) this quantum computation would take ~10^25 years for a classical computer to simulate, it would also take ~10^25 years for a classical computer to directly verify the quantum computer’s results!! (For example, by computing the “Linear Cross-Entropy” score of the outputs.) For this reason, all validation of Google’s new supremacy experiment is indirect, based on extrapolations from smaller circuits, ones for which a classical computer can feasibly check the results. To be clear, I personally see no reason to doubt those extrapolations. But for anyone who wonders why I’ve been obsessing for years about the need to design efficiently verifiable near-term quantum supremacy experiments: well, this is why! We’re now deeply into the unverifiable regime that I warned about.
Note that the OP wrote "you MUST compute something that is impossible to do with a traditional computer". I demonstrated a simple counter-example to this statement: you CAN demonstrate forward progress by factorizing big numbers, but the problem is that no one can do it despite billions of investments.
They claimed it last time in 2019 with Sycamore, which could perform in 200 seconds a calculation that Google claimed would take a classical supercomputer 10,000 years.
That was debunked when a team of scientists replicated the same thing on an ordinary computer in 15 hours with a large number of GPUs. Scott Aaronson said that on a supercomputer, the same technique would have solved the problem in seconds.[1]
So if they now come up with another problem which they say cannot even be verified by a classical computer and uses it to claim quantum advantage, then it is right to be suspicious of that claim.
1. https://www.science.org/content/article/ordinary-computers-c...
Yes, quantum supremacy on an artificial problem is quantum supremacy (even if it's "this quantum computer can simulate itself faster than a classical computer"). Quantum supremacy on problems that are easy to verify would of course be nicer, but unfortunately not all problems happen to have an easy verification.
In general we have thousands of optimisations problems that are hard to solve but immediate to verify.
What's factually wrong about it? OP said "you must compute something that is impossible to do with a traditional computer" which is true, regardless of the output produced. Verifying an output is very different from verifying the proper execution of a program. The difference between testing a program and seeing its code.
What is being computed is fundamentally different from classical computers, therefore the verification methods of proper adherence to instructions becomes increasingly complex.
The point stands that for actually interesting problems verifying correctness of the results is trivial. I don't know if "adherence to instructions" transudates at all to quantum computing.
My understanding is that many problems have solutions that are easier to verify than to solve using classical computing. e.g. prime factorization
Well, yes and no. This is only true because we are talking about closed models from closed companies like so-called "OpenAI".
But if all models were truly open, then we could simply verify what they had been trained on, and make experiments with models that we could be sure had never seen the dataset.
Decades ago Microsoft (in the words of Ballmer and Gates) famously accused open source of being a "cancer" because of the cascading nature of the GPL.
But it's the opposite. In software, and in knowledge in general, the true disease is secrecy.
How do you verify what a particular open model was trained on if you haven’t trained it yourself? Typically, for open models, you only get the architecture and the trained weights. How can you reliably verify what the model was trained on from this?
Even if they provide the training set (which is not typically the case), you still have to take their word for it—that’s not really "verification."
Lots of ai researchers have shown that you can both give credit and discredit "open models" when you are given a dataset and training steps.
Many lauded papers fell into reddit Ml or twitter ire when people couldnt reproduce the model or results.
If you are given the training set, the weights, the steps required, and enough compute, you can do it.
Having enough compute and people releasing the steps is the main impediment.
For my research I always release all of my code, and the order of execution steps, and of course the training set. I also give confidence intervals based on my runs so people can reproduce and see if we get similar intervals.
If not, it's not really "open", it's bs-open.
If they've done it right, you can re-run the training and get the same weights. And maybe you could spot-check parts of it without running the full training (e.g. if there are glitch tokens in the weights, you'd look for where they came from in the training data, and if they weren't there at all that would be a red flag). Is it possible to release the wrong training set (or the wrong instructions) and hope you don't get caught? Sure, but demanding that it be published and available to check raises the bar and makes it much more risky to cheat.
I wonder what the response of working mathematicians will be to this. If the proofs look credible it might be too tempting to try and validate them, but if there’s a deluge that could be a hug time sync. Imagine if Wiles or Perelman had produced a thousand different proofs for their respective problems.
It helps me a lot when I feel lost. It's often wrong in the calculations, but it's cool to have a study buddy that doesn't judge you.
If I get blocked with a problem I can't solve, I ask for assistance with my approach.
I enjoy asking ChatGPT about the context behind all that math theory. It's nice to elaborate on that as most of the math books are very lean and provide no applied context.
But the problem then is that one can suppose there are also true short statements in ZFC which likewise require doubly exponential time to reach via any path. Presburger Arithmetic is decidable whereas ZFC is not, so these statements would require the additional axioms of ZFC for shorter proofs, but I think it's safe to assume such statements exist.
Now let's suppose an AI model can resolve the truth of these short statements quickly. That means one of three things:
1) The AI model can discover doubly exponential length proof paths within the framework of ZFC.
2) There are certain short statements in the formal language of ZFC that the AI model cannot discover the truth of.
3) The AI model operates outside of ZFC to find the truth of statements in the framework of some other, potentially unknown formal system (and for arithmetical statements, the system must necessarily be sound).
How likely are each of these outcomes?
1) is not possible within any coherent, human-scale timeframe.
2) IMO is the most likely outcome, but then this means there are some really interesting things in mathematics that AI cannot discover. Perhaps the same set of things that humans find interesting. Once we have exhausted the theorems with short proofs in ZFC, there will still be an infinite number of short and interesting statements that we cannot resolve.
3) This would be the most bizarre outcome of all. If AI operates in a consistent way outside the framework of ZFC, then that would be equivalent to solving the halting problem for certain (infinite) sets of Turing machine configurations that ZFC cannot solve. That in itself itself isn't too strange (e.g., it might turn out that ZFC lacks an axiom necessary to prove something as simple as the Collatz conjecture), but what would be strange is that it could find these new formal systems efficiently. In other words, it would have discovered an algorithmic way to procure new axioms that lead to efficient proofs of true arithmetic statements. One could also view that as an efficient algorithm for computing BB(n), which obviously we think isn't possible. See Levin's papers on the feasibility of extending PA in a way that leads to quickly discovering more of the halting sequence.
This is a correct statement about the worst case runtime. What is interesting for practical applications is whether such statements are among those that you are practically interested in.
If the truth of these higher level statements instantly unlocks many other truths, then it makes sense to think of them in the same way that knowing BB(5) allows one to instantly classify any Turing machine configuration on the computation graph of all n ≤ 5 state Turing machines (on empty tape input) as halting/non-halting.
If every true theorem had a proof in a computationally bounded length the halting problem would be solvable. So the AI can't find some of those proofs.
The reason I say 3 is deep is that ultimately our foundational reasons to assume ZFC+the bits we need for logic come from philosohical groundings and not everyone accepts the same ones. Ultrafinitists and large cardinal theorists are both kinds of people I've met.
This has little to do with the usefulness of LLMs for research-level mathematics though. I do not think that anyone is hoping to get a decision procedure out of it, but rather something that would imitate human reasoning, which is heavily based on analogies ("we want to solve this problem, which shares some similarities with that other solved problem, can we apply the same proof strategy? if not, can we generalise the strategy so that it becomes applicable?").
Why do you say this? The AI doesn't know or care about soundness. Probably it has mathematical intuition that makes unsound assumptions, like human mathematicians do.
> How likely are each of these outcomes?
I think they'll all be true to a certain extent, just as they are for human mathematicians. There will probably be certain classes of extremely long proofs that the AI has no trouble discovering (because they have some kind of structure, just not structure that can be expressed in ZFC), certain truths that the AI makes an intuitive leap to despite not being able to prove them in ZFC (just as human mathematicians do), and certain short statements that the AI cannot prove one way or another (like Goldbach or twin primes or what have you, again, just as human mathematicians can't).
Society is CLEARLY not ready for what AI's impact is going to be. We've been through change before, but never at this scale and speed. I think Musk/Vivek's DOGE thing is important, our governent has gotten quite large and bureaucratic. But the clock has started on AI, and this is a social structural issue we've gotta figure out. Putting it off means we probably become subjects to a default set of rulers if not the shoggoth itself.
Previously workers in a field disrupted by automation would retrain to a different part of the economy.
If AI pans out to the point that there are mass layoffs in hundreds of sectors of the economy at once, then i’m not sure the process we have haphazardly set up now will work. People will have no idea where to go beyond manual labor. (But this will be difficult due to the obesity crisis - but maybe it will save lives in a weird way).
> programmers are still in the denial phase
I am doing a startup and would jump on any way to make the development or process more efficient. But the only thing LLMs are really good for are investor pitches.
But I am confident the answer to the question in the headline is "no, not for several decades." It's not just the underwhelming benchmark results discussed in the post, or the general concern about hard undergraduate math using different skillsets than ordinary research math. IMO the deeper problem still seems to be a basic gap where LLMs can seemingly do formal math at the level of a smart graduate student but fail at quantitative/geometric reasoning problems designed for fish. I suspect this holds for O3, based on one of the ARC problems it wasn't able to solve: https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_pr... (via https://www.interconnects.ai/p/openais-o3-the-2024-finale-of...) ANNs are simply not able to form abstractions, they can only imitate them via enormous amounts of data and compute. I would say there has been zero progress on "common sense" math in computers since the invention of Lisp: we are still faking it with expert systems, even if LLM expert systems are easier to build at scale with raw data.
It is the same old problem where an ANN can attain superhuman performance on level 1 of Breakout, but it has to be retrained for level 2. I am not convinced it makes sense to say AI can do math if AI doesn't understand what "four" means with the same depth as a rat, even if it can solve sophisticated modular arithmetic problems. In human terms, does it make sense to say a straightedge-and-compass AI understands Euclidean geometry if it's not capable of understanding the physical intuition behind Euclid's axioms? It makes more sense to say it's a brainless tool that helps with the tedium and drudgery of actually proving things in mathematics.
Which is actually a problem I have with ARC (and IQ tests more generally): it is computationally cheaper to go from ARC transformation rule -> ARC problem than it is the other way around. But this means it’s pretty easy to generate ARC problems with non-unique solutions.
The database stopped being secret when it was fed to proprietary LLMs running in the cloud. If anyone is not thinking that OpenAI has trained and tuned O3 on the "secret" problems people fed to GPT-4o, I have a bridge to sell you.
Edit: I do see from your profile that you are a real person though, so I say this with more respect.
Sam Altman has long talked about believing in the "move fast and break things" way of doing business. Which is just a nicer way of saying do whatever dodgy things you can get away with.
I understood the statements of all five questions. I could do the third one relatively quickly (I had seen the trick before that the function mapping a natural n to alpha^n was p-adically continuous in n iff the p-adic valuation of alpha-1 was positive)
Seems to me the answer to 'Can AI do maths yet?' depends on what you call AI and what you call maths. Our old departmental VAX running at a handfull of megahertz could do some very clever symbol manipulation on binomials and if you gave it a few seconds, it could even do something like theorum proving via proto-prolog. Neither are anywhere close to the glorious GAI future we hope to sell to industry and government, but it seems worth considering how they're different, why they worked, and whether there's room for some hybrid approach. Do LLMs need to know how to do math if they know how to write Prolog or Coc statements that can do interesting things?
I've heard people say they want to build software that emulates (simulates?) how humans do arithmetic, but ask a human to add anything bigger than two digit numbers and the first thing they do is reach for a calculator.
You won't get the Turing machine evaluation mechanism and determinism, but you will have a generator. Although the viability of what is generated is is question. Because the other part of formalism, semantics, is almost always missing.
So the higher the cost the better the performance. While models and hardware can be improved the curve is still steep.
The big answer is what are people using it for? We'll they are using lightweight simplistic models to do targeted tasks. To do many smaller and easier to process tasks.
Most of the news on AI is just there to promote a product to earn more cash.
If that's referring to Large Language Models, meaning everything after the fist GPT and BERT, then that's absolutely not right. The first LLM that demonstrated the ability to generate coherent, fluently grammatical English was GPT-2. That story about the unicorns- that was the first time a statistical language model was able to generate text that stayed on the subject over a long distance and made (some) sense.
GPT-2 was followed by GPT 3 and GPT 3.5 that turned the hype dial up to 11 and were certainly "public" at least if that means publicly available. They were coherent enough that many people predicted all sorts of fancy things, like the end of programming jobs and the end of journalist jobs and so on.
So, weird statement that one and it kind of makes me wary of Gell-Mann amnesia while reading the article.
Now, all of that will be done by AI.
Reminds of the time when I finally enabled invincibility in Goldeneye 007. Rather boring.
I think we've stopped to appreciate the human struggle and experience and have placed all the value on the end product, and that's we're developing AI so much.
Yeah, there is the possibility of working with an AI but at that point, what is the point? Seems rather pointless to me in an art like mathematics.
No "AI" of any description is doing novel proofs at the moment. Not o3, or anything else.
LLMs are good for chatting about basic intuition with, up to and including complex subjects, if and only if there are publically available data on the topic which have been fed to the LLM during its training. They're good at doing summaries and overviews of specific things (if you push them around and insist they don't waffle and ignore garbage carefully and keep your critical thinking hat on, etc etc).
It's like having a magnifying glass that focuses in on the small little maths question you might have, without you having to sift through ten blogs or videos or whatever.
That's hardly going to replace graduate students doing proofs with professors, though, at least not with the methods being employed thus far!
So, for any LLM, if you intersperse more than that number of ‘X’ tokens between each useful token, they won’t be able to do anything resembling intelligence.
The current LLMs are a bit like n-gram databases that do not use letters, but larger units.
Naturally, humans couldn’t do it, even though they could edit the input to remove the X’s, but shouldn’t we evaluate the ability (even intelligent ability) of LLM’s on what they can generally do rather than amplify their weakness?
I am not claiming LLMs aren’t or cannot be intelligent, not even that they cannot do magical things; I just rebuked a statement about the lack of limits of LLMs.
> Naturally, humans couldn’t do it, even though they could edit the input to remove the X’s
So, what are you claiming: that they cannot or that they can? I think most people can and many would. Confronted with a file containing millions of X’s, many humans will wonder whether there’s something else than X’s in the file, do a ‘replace all’, discover the question hidden in that sea of X’s, and answer it.
There even are simple files where most humans would easily spot things without having to think of removing those X's. Consider a file
How X X X X X X
many X X X X X X
days X X X X X X
are X X X X X X
there X X X X X X
in X X X X X X
a X X X X X X
week? X X X X X X
with a million X’s on the end of each line. Spotting the question in that is easy for humans, but impossible for the current bunch of LLMsHmm, I wonder if adding a compression layer during encoding helps?
Humans have weak attention compared to it, this is a poor example.
Even if that weren't true, the context windows can be quite large and will get bigger as people figure out how to optimize LLMs. For example Gemini 1.5 has a context of 2 million tokens. A book is typically around 120,000 words, so that's almost 20 books. So one could argue with a context this big they could construct reasoning chains involving far more disparate pieces of information than humans typically work with simultaneously, and arguably that demonstrates intelligence as well.
the management class has a strong incentive to believe in this narrative, since it helps them reduce labor cost. so they are investing in it.
eventually, the emperor will be seen to have no clothes at least in some usecases for which it is being peddled right now.
Money cannot solve the issues faced by the industry which mainly revolves around lack of training data.
They already used the entirety of the internet, all available video, audio and books and they are now dealing with the fact that most content online is now generated by these models, thus making it useless as training data.
This might be a dogma that needs to die.
Do you? Anything anywhere you could point me to?
The algorithms live entirely off the training data. They consistently fail to "abduct" (inference) beyond any language-in/of-the-training-specific information.
Anyhow, philosophically speaking you are also only exposed to what your senses pick up, but presumably you are able to infer things?
As written: this is a dogma that stems from a limited understanding of what algorithmic processes are and the insistence that emergence can not happen from algorithmic systems.
AI doesn't need to learn everything, our LLM Models already contain EVERYTHING. Including ways of how to find a solution step by step.
Which means, you can tell an LLM to translate whatever you want, into a logical language and use an external logic verifier. The only thing a LLM or AI needs to 'understand' at this point is to make sure that the statistical translation from left to right is high enough.
Your brain doesn't just do logic out of the box, You conclude things and formulate them.
And plenty of companies work on this. Its the same with programming, if you are able to write code and execute it, you execute it until the compiler errors are gone. Now your LLM can write valid code out of the box. Let the LLM write unit tests, now it can verify itself.
Claude for example offers you, out of the box, to write a validation script. You can give claude back the output of the script claude suggested to you.
Don't underestimate LLMs
The answer is yes, it can utilize a stateful python environment and solve complex mathematical equations with ease.
https://chatgpt.com/share/676980cb-d77c-8011-b469-4853647f98...
More advanced solutions:
https://chatgpt.com/share/6769895d-7ef8-8011-8171-6e84f33103...
I’ve yet to encounter an equation that 4o couldn’t answer in 1-2 prompts unless it timed out. Even then it can provide the solution in a Jupyter notebook that can be run locally.
Right, so that reasoning is based on what, exactly?
I've yet to encounter one of such sessions where the person is not handholding the LLMs during the whole process, basically describing the solution (instead of the problem) in natural language.
I really really wish AI would make some breakthrough and be really useful, but I am so skeptical and negative about it.
The "open worm project" was an effort years ago to get computer scientists involved in trying to understand what "software" a very small actual brain could run. I believe progress here has been very slow and that an idea of ignorance that much larger brains involve.
Why?
What does it mean to give a model a calculator?
What do you mean “let it show its working”? If I ask an LLM to do a calculation, I never said it can’t express the answer to me in long-form text or with intermediate steps.
If I ask a human to do a calculation that they can’t reliably do in their head, they are intelligent enough to know that they should use a pen and paper without needing my preemptive permission.
Very many things conventionally labelled in the 50's.
You are speaking of LLMs.
Which is, wrongly: so, don't spread the bad notion and habit.
Bad notion and habit which has a counter-helpful impact on debate.
But maths are also fun and fulfilling activity. Very often, when we learn a math theory, it's because we want to understand and gain intuition on the concepts, or we want to solve a puzzle (for which we can already look up the solution). Maybe it's similar to chess. We didn't develop search engines to replace human players and make them play together, but they helped us become better chess players or understanding the game better.
So the recent progress is impressive, but I still don't see how we'll use this tech practically and what impacts it can have and in which fields.