If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
9,765 karma · joined February 4, 2012
If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
I explicitly give credit to the papers the proof builds upon. Quoting from the article:
>The proof was only possible thanks to the many existing results from [References](https://gaearon.github.io/conway-refinement/#/references). In particular, [A factorisation theory for generalised power series and omnific integers](https://doi.org/10.1016/j.aim.2024.109513) by S. L’Innocente and V. Mantova has played a crucial role in the proof.
The project README (https://github.com/gaearon/conway-refinement) also states:
>Over the years, Berarducci, Pitteloud, Pommersheim and Shahriari, and L'Innocente and Mantova developed increasingly strong results about exactly that. This proof builds on all of that work and claims to finish the refinement conjecture.
It includes the same list of references: https://github.com/gaearon/conway-refinement#references
Regarding who I've talked to over email, I wasn't sure I want to "out" them explicitly in the article due to the article's likely polarizing nature, but, since he already engaged in this thread, I can confirm that this person was Prof. Mantova (who is a co-author of the primary paper that the result I published rests upon, and who I did explicitly credit in the text of the article and in the references).
You can read Prof. Mantova's own reflection here: https://news.ycombinator.com/item?id=49761718
I have not had an extended conversation with anyone else.
>Following your analogy it is not a very fair game if it allows winning for those having access to latest exploits. Playing against skiddies who bought the hack kind of ruins it.
You misunderstood the point of my analogy. The point was not that it is a game in which you "win" or "lose" and thus I have "won" or something like that.
The point was that "obtaining a true proof without understanding it" doesn't imply "understanding has no value" any more than "completing a game in an unintended way" implies "playing has no value". You stated a false implication. That false implication is what I am highlighting with the analogy. Understanding has a societal value; doing something weird for the fun of it has value to the person doing it; sometimes people do things with no value at all, and that's okay too. Each individual act is not a statement about what kinds of things have value.
Personally, I do not see mathematics as a game of "win" and "lose". If the broad value is primarily in understanding (which I think is correct), then my "contribution" doesn't make much dent in that. However, it might, as Prof. Mantova mentions, help show the way to an actually useful proof that improves collective understanding. That would please me beyond merely having had fun. In any case, I've enjoyed the experience of it in part due to the absurdity of it. Maybe it requires a certain kind of sense of humor.
Regarding the models, as far as I know, all of the models I have used are currently broadly available. Yes, they require subscriptions so access is not democratized yet. There's nothing I can do about that. I think it's unreasonable to require that someone playing around with LLM's must use the inferior free models.
Re: "rich", I've essentially spent $400 on this (in subsidized subscriptions), plus my free time being a mindless drone. Given that you assume my role is automatable, it sounds like this is relatively accessible to anyone with $400 (as long as AI companies continue subsidizing the frontier models). I don't think I've had some kind of an unfair advantage beyond that. If anything, a proper mathematician would probably be able to derive the result much faster with the same tools.
I don't know how the broad availability of these tools (to mathematicians and non-mathematicians alike) will change the field, what is considered prestigious, what work gets funding, how it affects the pipeline, etc. You seem to be implying that even testing the limits of these tools, or at least publishing the results obtained with them, is unethical in itself, even though it is broadly accessible now. I can understand this point of view.
I am not saying that my work constitutes a mathematical contribution on its own. Not any more than stumbling upon an anonymous manuscript with the solution would constitute a mathematical contribution. I do, however, think that it can lead to a mathematical contribution if any mathematicians consider it worthwhile to do something with it. Whether or not they consider it worthwhile is not up to me.
It might be possible to plainly continue-shot it with more powerful models in the future. I agree that whatever I did is probably automatable.
The alien in the analogy is the corpus of knowledge that’s newly reachable via LLMs.
Would you, in this situation, be mad at the aliens? Would you say the aliens have "extracted value"? This is kind of how I see this project.
I do that too when working on software. Here, I did a little bit of that, but it is much harder when I have almost no domain knowledge (aside from understanding the statement of the conjecture), since at each point the LLM might trick me anyway, and it would probably take me a year to understand the concepts enough to tell when what it's saying doesn't make sense.
>Even just keeping the Lean more up to date would have likely saved tokens
That's the conclusion I came to by the final week! But don't underestimate how much Lean needed to be written in the first place to even "catch up" with the reference papers. I had no option to keep it up to date when I started.
To answer directly: I think understanding is the primary value. And I have reasons to hope that my "work" here can ultimately contribute to a better understanding: https://news.ycombinator.com/item?id=49761718. However, there are also other reasons to do things than producing value. You can do things for fun, to show they're possible, to ask meta questions about the field.
Regarding calling the proof "mine", I've replied to that here: https://news.ycombinator.com/item?id=49755885.
Yes, I was not doing any mathematical work in curating the output, but the article makes it quite clear that pivots and constraints the LLM would not impose on itself were critical to actually making progress.
Also:
>He doesn’t even know any mathematics and has never cared to learn
While I don't know enough mathematics to work on this problem, claiming something like this is preposterous. As I link in the first paragraph of the article, I've been learning mathematics on my own by going through Terence Tao's Analysis book and solving exercises. I'm familiar with the concepts of mathematical definitions, proofs, etc. I've gotten about halfway through the book solving them on paper before abandoning it (and later got through the first few chapters in Lean, also solving every exercise — by hand, mind you). Sure, this doesn't make me a mathematician, but I'm closer to a dropout first-year student than to someone who has "never cared to learn".
The happy case here is that my obtuse proof leads to a concise and illuminating mathematical proof, which is exactly my hope for the endeavour. I think this could then be a positive example of AI/human and amateur/professional collaboration.
What would a positive example look like to you? What do you think the manifesto argues for?
I’m slightly stretching the metaphor here because the number lines gives me enough structure (order expressed visually) and I only deal with at most one set item at a time (since we construct in the order of simplicity and can use the already constructed numbers), so it (IMO) unnecessarily complicates things to even talk about sets when we’re sort of just making cuts on the line. But in either case I don’t see the problem with colloquially saying “nothing” here.
For the infinitesimal number, I think it makes more sense to use {0}|{1,1/2,1/4,1/8,...} since it gets born at the same day as say 1/3. So it is easier to understand how it arises without "waiting" for all reals.
> We may say that Cantor was only interested in moving ever rightwards, whereas Dedekind stopped to fill in the gaps, so that R was always empty for Cantor, never empty for Dedekind. It is remarkable that by dropping these restrictions we obtain a theory that is both more general and more easy to work with.
This is precisely the intuition I present to the reader of the article. I am relying on visual aid (concretely, the ordered number line) to imply the machinery explicit in the actual recursive definition. The intended reader of this article is not a mathematician, and I think intuition is vastly more important here.
And I don't think I'm conflating 0 with {0} as you claim. When I say zero is "between nothing and nothing", I mean 0 := {|}. When I say one is "between zero and nothing", I mean 1 := {0|}. When I say 1/2 is "between 0 and 1", I mean 1/2 := {0|1}. And so on. I elide "the simplest number" because I am already going in the order of simplicity. I do not need to explain that alternative spellings like 1/2 = {0.2 | 1} are valid because it is not relevant to establishing the mental model of birthdays.
For the finite cases in my explanation, I do not need to explain that the left and the right parts form sets because I only ever need at most one surreal on either side to define the next generation. I also do not need to state the left/right order condition because it is already visually implied by the picture. For the same reason, I do not need to explicitly quantify over the set of earlier-born surreals, since in these finite cases, if we go birthday by birthday, each next day's surreals are definable via the numbers already constructed by the previous day.
I agree that these finite examples don't spell out how to handle infinitely many bounds at the omega-th day, which is where I believe the illustration embedded below is more helpful. I still think "a gap beyond 0, 1, 2, 3, ... with nothing on the right" is a useful intuition when we get there.
For a more precise but accessible treatment, I think https://www.infinitelymore.xyz/p/surreal-numbers is much clearer than Wikipedia.
Primarily I thought of this as a sort of "epistemic performance art project", maybe similar to playing Elden Ring blindfolded having never played it before, or speedrunning a game by opening a box a thousand times and overflowing some counter. It's funny and absurd to do knowledge work without the knowledge.
I think it's also a stress test of meta skills. Like, how much can we do without knowing? What kind of processes can we set up around these demons that would constrain them into our requirements? How can we know when things are going wrong? In some sense, this isn't too different from engineering management.
Naturally, I'm also interested in how much of my role in this could've been automated away. Can there be a skill for that? Then "do a breakthrough" is an irrelevant implementation detail of that skill.
Note that "do a breakthrough" actually produced the worst results over the runs. The best results were from more directed runs like searching for first obstacle towards the next milestone.
Let me first clarify my relationship with mathematics. I think of myself as "an awestruck observer from a distance". I find some parts that I understand beautiful, and I have also tried to understand some of the basics rigorously. However, I generally just can't make my way through any serious paper, as I both lack the prerequisites and struggle with the amount of inference mathematics tends to place on the reader. That's the "from a distance" part.
Now, about picking the problem. I am genuinely "pulled by" surreal numbers themselves. I find them irresistibly beautiful. There is also a bit of bitterness around how they haven't fulfilled their promise (yet?) as Conway hoped they would be able to become a better foundation for some mathematics. But they are a bit too difficult to prove things about so far, and we know too little about them. So what "pulls me" also is a possibility of making enough dents in this that we would be able to use them more broadly, and learn even more things about them.
However, I do not know the details of the latest research. I don't know which problems have actually been solved, which pursue Conway's original vision vs narrower approaches, and which are elegant enough to feel "awestruck" enough about. So this is an invitation from me to LLM to share what it "feels pulled by" (for whatever definition; I think of it as just navigating the languagespace) , and then sifting through that list to see if something it lists makes me feel something. I would assume that with the field currently being so small (serious mathematicians mostly don't care about surreals), it's easy to get the LLM "excited" (again, just a vector in the languagespace) enough that it would give me genuinely interesting candidates. Then it's up to me to sift through them and see if they "speak" to me.
It's like asking a mathrock nerd to share their favorite mathrock albums. Niche enough that you'd likely get good results. Then you can listen and form an opinion.
In this particular example, the "ONAG birthday" and "maybe last Conway's unsolved conjecture about surreals" part spoke to me emotionally, the statement itself amazed me with its simplicity, and I felt "blood in the water" related to the recent results bringing the conjecture closer. So I felt the pull myself and went with it.
While Lean is tightening things up after the recent LLM-driven hacks, I agree that bugs are possible. Although usually code that exploits them is obviously aggressive and is deliberately using the more obscure features related to metaprogramming. Also note that my solution has passed the nanoda kernel as well (https://palomar-registry.org/entry?id=PALOMAR-2026-09-03-000...).
That said, again, I never implied that I'm asking mathematicians to "laboriously [check] the generated proof" which is what your parent comment says. The value to mathematicians is knowing that the conjecture is probably right, and knowing the rough path the LLM has taken to it. Instead of checking the Lean proof line by line, what mathematicians are interested in doing (at least, the ones I've been in contact with) is finding a shorter and more direct proof now that they're aware of the outline and main intermediate claims. As for how much value they find in that, I presume they would be able to speak to that when/if they would like to make their research public.