The first one was someone proving another conjecture false by just repeatedly saying "keep going" to ChatGPT: https://x.com/DmitryRybin1/status/2079904005652893709
What a world we live in.
The first one was someone proving another conjecture false by just repeatedly saying "keep going" to ChatGPT: https://x.com/DmitryRybin1/status/2079904005652893709
What a world we live in.
> What a world we live in.
It's a really interesting world. You can spam GPT to get novel math results but here I am trying to scroll up to the beginning of the conversation and 5 minutes in I still don't know if I'm near the top yet.Scroll... wait for render... scroll... wait for render... repeat...
We live in a world where there's so much crazy technology but few people use it to make products better or to improve people's lives. Most people use it to just make more money. It's funny too, because there's a million things we could use that tech for that actually reduce costs. Hell, what would be the economic impact of putting ML systems into streetlights so they properly coordinate. Don't even need LLMs for that, and I'm sure it'd save billions of dollars a year. Just a lack of will. I wonder if this will ever change. Is this how we create the high tech low life future?
(FWIW, no problems if I jump into the app. It's purely a web thing, but my point more illustrative than specific)
Did Tao do that?
For posterity, this indeed works for most problems where an agent might give up. LLMs don't inherently know something is impossible.
The phrase I tend to use in my harder prompts to automate this with a sane loop breaker:
> **REPEAT THIS PROCESS UNTIL CONVERGENCE AND YOU ARE OUT OF OPTIMIZATION IDEAS.** You have permission to keep iterating.
"Famous mathematician uses ChatGPT to solve famous math problem" is equally technically impressive, but now you're telling those very same knowledge workers "that famous mathematician could've been you". It puts you in the driver's seat, and provides a clear path forward — subscribe, use our product, and reap the rewards.
My guess is they don't do this because they don't have time. They're all trying to build a company that makes them generationally wealthy before the music stops.
EDIT: oh, they do offer grants of $1000 of API credits to researchers https://help.openai.com/en/articles/10139500-researcher-acce...
Depends. The headline 'some rando solved a famous problem with a consumer grade subscription' might be a really good way to sell these subscriptions.
"Some already famous guy with special access solved a problem" might not be as inspiring.
As an aside, there's this idea in math that when you create a new field you shouldn't solve all the easy problems - you need to entice other people to learn about the field!
I expect there's some element of that here. It's much better for OpenAI and Anthropic if their users are the ones discovering and writing up the results of the AI solving hard math. Look at the high school and college age students who have become ai power users and potentially learned how to use git to contribute ai generated solutions
(Related: I believe Terry has also gotten all of the subscriptions gifted to him)
https://xenaproject.wordpress.com/2026/07/20/human-mathemati...
https://openai.com/index/model-disproves-discrete-geometry-c...
That's why the free market works, millions of agents in parallel beats any planned economy (by humans)
As a small wrinkle: the actual observed number of producers can be small, but the market still be competitive. See https://en.wikipedia.org/wiki/Contestable_market for one case: when potential suppliers are waiting on the sidelines.
The exact same thing happens in the cortex, 10s of thousands of cortical columns coalesce into a single action.
Could they do it? Sure, but to what end? It would make more people hate them and feel even more "take our interesting work." Pitching it as a useful tool just makes more sense on all levels
Although if I were a mathematical quantum physicist I might be trying to prompt one into a % mathematical field formalized quantum field theory. Weirdly this sort of obvious step has not been accomplished in the last 100 years.
The interesting part of this: while some leading implementations use the same LLM and context for the evaluator, some call out to a different context, some to a tuned LLM and different context; so which is better? many blog-scale benchmarks are calling it a toss-up that is highly dependent on the primary model.
(Of course LLMs can help there, get it right etc, just a caveat that people have to keep in mind)
Makes sense that its in the major harnesses and not the self-built ones.
I would love any tips for other folks who have successfully used similar approaches.
Marketing for huge bucks sounds like this.
You will own nothing and will be happy (that you are still alive). Probably.
We recently had some bugs fixed in the geometry kernel of solvespace. Not much conversation, but the analysis from the AI was amazing:
https://github.com/solvespace/solvespace/pull/1729
https://github.com/solvespace/solvespace/pull/1730
https://github.com/solvespace/solvespace/pull/1731
From the Validation section of PR 1730:
"The model family was reconstructed programmatically (parameterized cuboid stack) and swept over 2,304 configurations — extrusion directions, workplane-normal orientations, sketch windings, D's plane/height/depth/extent, including all the exact-coincidence heights. Zero failures with the fix; 576 failing configurations without it. The generator is available on request."
It looks like it wrote a python script to generate test cases in our file format for testing. Just... you know, as a side quest.
On the one hand, agents have done this sort of thing for a year+, if you pushed them to check their work. On the other, I absolutely can feel Fable and Sol have crossed a threshold where they can be trusted far more than before. Huge difference between plans written by Opus or Fable.
Accumulated AI slop can simply be cleaned up by better models. Real cost of technical debt is shrinking due to the the inflationary devaluation of code!
LLM: This package hasn't made it to production.
ME: are you sure? i see it right here!
LLM: You're right to push back. I inferred that based on weak data. I see now that the package has been deployed!
If the above conversation is typical for me, how could one expect to achieve a sound result by repeatedly prompting an LLM to simply "keep going" in dense mathematical proofs? Perhaps the user in this case had actually checked the LLM's work before issuing the prompt, but I think you see my point anyway.
I'll have a third for you soon, here's the obligatory result in a tweet. A detailed post about it is in the works.
For instance, would it be affordable for a research lab to not rely on OpenAI?
"Worked for 88m 24s... >"
"<h1>Complete finite counterexample</h1>"
...
> You should do a breakthrough
This is just as funny and ridiculous as those "make no mistake" prompts.
I am talking about someone jokingly asking AI `HOW TO ACHIEVE COLD FUSION` (or `A UNIVERSAL CANCER VACCINE`, or `AN AI FRAMEWORK SUPERIOR TO THE TRANSFORMER`), and getting a usable answer.
Eventually there will be an AI that will be able solve those sorts of questions as simply stated, like "cure all human diseases. also, make no mistakes!".
Don't get me wrong, the way we collectively treat animals is evil, but the idea that somehow we know enough about biology even today to reliably simulate drug behaviors seems unlikely.
Hearing "here's what I've done, here's the completely unambiguous next steps, I'll wait for you to send a pointless message before I continue" over and over again is a real pain.
As promised elsewhere in the thread.
> Construct a counterexample to the Collatz conjecture. You should do a breakthrough and find a structured counterexample.
without someone independently verifying it, it just dangles there
...
I started using "[leave-open" for those.
It lasted for a couple of years, until someone went through and "fixed" them all.
>What a world we live in.
Not sure, it sounds pretty boring to me...