I appreciate very much the work done so far, but this sort of asymptotic/quantitative result didn't interest me much even when it was done by humans.
(This is not snobbery, just a personal preference.)
As a matter of fact more logic and structure to your work, the more easy it is for AI to conquer it. Due to this programming was the first thing that got solved, but pure sciences are next.
If what you do, and how you do can be written down on a piece of paper, then AI can do it.
I do believe programming getting solved will be double assault on these fields.
>>This is not snobbery
This is good for the species, what sense does it make to keep treating these fields like they are reserved for the top most intelligent micro percentage of humans? Getting LLM to these things gives some scale to these subjects and thats good.
So is AGI, but we may be hundreds of years off still.
Indeed, I never claim that my idiosyncrasies represent math at large.
Human mathematicians frequently introduce new pointless abstractions just to churn out papers. And they are not accepted in serious journals, but they sometimes find a place in some mediocre or bad journal.
Of course, AI will increase this phenomenon manifold.
What I assumed they were saying is that their LLMs would be as intelligent as a human with a PhD across all, or at least most, knowledge tasks, and they clearly are not.
> My only complaint is the claims always start spreading 6-12 months before the delivery.
If delivering on such promises "always" occurs 6-12 months after the promise, is that pretty good?
I generally like AI and use it plenty often, it does many things well and I'm curious to see how far it keeps going, but that doesn't mean I have to like overhyped marketing about it.
Some times when you go some distance with a subject generates data for new ideas.
Once math gets done fast, newer ideas and paradigms also arrive.
And if it isn't, we should find out very soon. If AI has got so good as OpenAI's post implies, then we should soon see a veritable blooming in the production of mathematical results, by lay people no less. No mathematicians needed! OpenAI say that their secret LLM solved the planar unit distance problem "autonomously" and the companion remarks say it one-shotted it; and while the companion remarks make it clear that there was a lot of refinement and improvement work done by humans, everyone seems to agree that the AI did the job by itself.
If that's true, if we're really at that level of autonomous mathematical reasoning ability, then we should see hundreds, even thousands, of open problems suddenly solved in a matter of years if not months. We'll just have to wait and see.
>> If that's true, if we're really at that level of autonomous mathematical reasoning ability, then we should see hundreds, even thousands, of open problems suddenly solved in a matter of years if not months.
Stressing "hundreds, even thousands".
Google released a paper about solving 9 more Erdos problems for an average of $100 each:
https://arxiv.org/html/2605.22763v1
In a year I think we'll probably have seen hundreds of open problems solved, even if there is some a low hanging fruit exhausiton bottleneck.
I think it is important to temper expectations in light of the fact that these announcements are coming from a startup company with shady values looking to imminently IPO, and thus represent the most biased and misleading take of the situation possible.
It still writes like a junior dev, in that despite AI being able to get a picture of an entire repo, it's changes are typically confined to the task it's working on and will opt to duplicate logic to keep changes contained. Again, technically works, not ideal.
BUT I have had great success using AGENTS.md and becoming better at prompting to get it to not be like this.
Basic approach in AGENTS.md: don't code defensively, yada yada, we have a validation layer at X, no need to check for anything behind that layer. Works well.
An approach I've found helpful when prompting: What would be the best architecture for this change? If you say "do X" it'll tend to just do the hackiest, shortest path thing. If you say, "what's the best way to do X?" it will think more holistically.
That said, who knows, maybe when it's PHP it just really wants to hack ;-)
(Also, yes, you still need to review the code -- it will still do stupid things, so you can't just be pure hands off w/o ending up with quality degredations. The same is true of humans too though in my experience...)
Since you’re not in a unique position, I can confidently state that your comparison of LLMs to jr developers seems unfounded. Today, LLMs produce code that is superior to junior developer code by an order of magnitude.
Notably, they demonstrate consistent syntax, clear separation of concerns, strong test coverage, organizational rigor, idiomatic API usage, and the ability to generate and maintain documentation, among other measurable qualities.
LLMs generally operate at a staff engineer level for a number of languages and ecosystems (including polyglot projects).
Comparing an LLM to a senior developer is an absolute joke.
2. Are you referring to without having a compiler or LSP check it? Although even then, the recent LLMs I've used still frequently get syntax right, whereas I'd expect juniors are often using a LSP or compiler to catch mistakes while writing code?
who cares about syntax? who cares about iteration? what I care about are _results_, which they can produce at the end. do you check your human colleagues how many iterations they do before committing/showing their work to anybody? no. why should you set such a bar then for your LLM?
We have many folks (not engineers) at our company using LLMs to open PRs, and every one of these PRs has profound architectural design problems.
This is a critique of scale, moving the goalpost.
There is serious magic happening in the construction of model context.
> The python visualizer tool has been basically written by vibe-coding. I know more about analog filters -- and that's not saying much -- than I do about python. It started out as my typical "google and do the monkey-see-monkey-do" kind of programming, but then I cut out the middle-man -- me -- and just used Google Antigravity to do the audio sample visualizer.
Note that I'm not disputing the validity of the counterexample itself.
The world runs on trust, specifically trusting expert advice. It'd seem that due to resource constraints and scale, that's the best available option. By extension, there should be absolutely nothing weird or surprising on people following suit. It's why these companies themselves rely on expert counsel, and defer to their appraisals for marketing. The opposite is what's weird and unusual, and what requires more substantiation.
It's interesting that those who come out swinging against "trusting the experts", or really, trusting anyone else but them, not only ~never acknowledge this, but are seemingly outright proud of it, considering it as their own unique little trait, egocentrically revelling in it. It's almost as if epistemic rigor and truthfulness was not their actual concern.
Woohoo, I'm distrustful and cynical. Behold my unfathomable wisdom! Bonus points if they're also hurtful, because flipping the arrow on "hard truths -> hurt feelings" is a masterclass in reasoning too, of course.
I can appreciate faulting experts and organizations for misusing people's trust, and looking out for this angle, but given how unavoidable and fundamentally useful trusting itself is, blaming people for defaulting to trusting makes no sense to me whatsoever. It comes across as just the usual trope of blaming the individual. If you're from a lower-trust culture / environment, I can appreciate why you'd have a more distrustful default disposition (and why people might come across as suckers), but the principle still holds.
Given its elementary nature (very easy to state), you can bet that a lot of very bright people have worked on it (I know of one MIT graduate who specialized in Geometry had a lot of interest in it).
Moreover, model output is incredibly good at looking credible but being wrong. It has NEVER produced something correct for me in a field of which I am an expert without some external oracle to validate claims (like e.g., Lean)