130 karma · joined June 14, 2025
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
Also, it still seems that AI has a much different style from humans, with more brute force and using obscure literature results, and the future might still end up human/AI complementary. We aren't in an AlphaZero situation where the AI learns everything through self-play. (Yet? But we don't even seem to be moving that way much? Can anybody qualified help out?) Things are just moving really fast now and it's hard to process everything.
In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.
Math is incredibly rich, and even the simplest things have insanely complicated structure when you zoom in. However this all ends up, math is bigger than LLMs, and the people who claim it is getting "solved" and we are running out of open problems haven't stared into the abyss enough.
One thing I noticed was that more than a few were seriously sick in childhood and had to be homeschooled. This includes Edward Morley, Peter Higgs, René Descartes (though I'm not sure how rare it was at his time), and the mathematician Julia Robinson, who was bedridden with scarlet fever at 9 years old, then had to get tutoring to catch back up, and had this to say about it [2]:
> I have since read that a solitary childhood or, what amounts to the same thing, a period of isolation resulting from an illness is frequently noted in the early lives of scientists. I am not sure what the significance of this finding is. Obviously I had to amuse myself for long periods of time, but I didn’t do so with mathematics. I am inclined to think that what I learned during that year in bed was patience.
> By the time I was well enough to go back to school, I had missed more than two years. My parents arranged to have me tutored by a retired elementary school teacher. In one year, working three mornings a week, she and I went through the state syllabuses for the fifth, sixth, seventh, and eighth grades. It makes me wonder how much time must be wasted in classrooms.
Sidenote: I found the book [1] through asking a free LLM what source this quote might be referring to. They are reasonably good at this kind of literature search, especially because it's easy to judge whether they gave you something useful.
[1] https://archive.org/details/cradlesofeminenc0000goer_l9f8/pa...
[2] https://web.archive.org/web/20181207045746/https://www.maa.o...
Conversations with strangers can be hard to get going, but they aren't this bad.
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
[0]: https://arxiv.org/pdf/2503.23674 (now published at https://www.pnas.org/doi/epdf/10.1073/pnas.2524472123). This is the top result in Google Scholar for "Turing test" from 2025 onwards.
[2]: "Dull Rigid Human meets Ace Mechanical Translator" (https://www.cambridge.org/core/books/abs/once-and-future-tur... or alternative access methods thereof)
The problem is that as a lay user of Firefox, I don't know how you could even make it better in terms of features (I could see it being faster etc.).
"But in a regular biology lab, this isn’t the finished paper. It is the result you show at lab meeting and say: “This looks interesting. Now we need to figure out what the hell it does.”"
This is not how people write!
https://en.wikipedia.org/wiki/General_relativity_priority_di...
IMO, this is why we actually need mathematics -- as a field in which to learn what it means to know what you're talking about.
Submitted as its own entry, hope you don't mind: https://news.ycombinator.com/item?id=49774097
The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
The weak point in this is: how do you evaluate if a partial result is promising? If this cost ~$10M as suggested elsewhere in the thread, probably not even OpenAI can just throw that at everything?