jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.
If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.
"there is a nand gate A connected to gate B through these wires, connected to another gate C, etc. what is the output value of gate C if i place a 1 at this gate"
Sort of like doing math "in their heads" (i.e. your "when adding numbers"), they would get it very right for simple/small cases (although the answer could have been in their training), then some LLMs would get it for harder cases (which were clearly not in their training), and all would fail at some point. This was all without any tool calling.
After a year of thinking about it, I made an eval [0] with ever-complexifying nand circuits - like, truly, bananas circuits [1] - and some models, do, in effect (through chain of thought? mostly?), get the right answer. ((what's nice is that you can always make a circuit at the very edge of what all models can correctly solve))
Tool-calling 100000% solves this problem for sure (evaluating a nand gate is trivial). But if you even prompt an llm to do math like a 5/6th grader (i.e. do it digit by digit, carry the 1, etc.) - I am quite certain most llms can, in fact, add numbers.
But yeah. These piles of weights are fascinating in how they seem flawed one day ("how many r's") and magical at once.
[0] https://lockstep.greg.technology
[1] https://lockstep.greg.technology/c/?id=rand_s4161_g160_d8
Decades before the modern AI push I was marveling at the distinction between the sheer overwhelming computational power of the human brain, if measured from the perspective of how much math it is doing under the hood, and its utter ineptness at basic arithmetic compared to the tools we can build. The earliest, klunkiest, most garbage mechanical adding machines we ever produced, long before we improved them by literally over a dozen orders of magnitude, were still already way better than we are at basic arithmetic.
There is something profound I still have not fully put my finger on in how basic arithmetic is so easy for a machine, yet the decisions we routinely make with our neural nets has been the laborious effort of decades with us still not arriving yet even with the trillions now poured into AI for machines. And vice versa. Even that practice I alluded to that allows you to train yourself to do this task on demand easily would incorporate mathematical advancements in representations that took our species thousands of years to come up with, rather than being something you get "for free" just for being smart. Our brains casually run an entire human body through an unbelievably complicated external universe, yet struggle with basic arithmetic.
There's so many places where we have one architecture that's a bit better than another at one thing, and a bit worse at another, but in the end they can both do the job. Like, if we had to do all our programming in immutable languages and run all our imperative code through an O(n log n) worst-case penalty for immutable languages emulating imperative RAM, we'd survive just fine over all. But between neural architectures and conventional arithmetic on dedicated silicon is this dozen+ order of magnitude difference on tasks. It's a pretty wild disparity. Some of the reason is somewhat obvious, I don't want to make it sound like I'm completely mystified... I just think there's probably, somewhere, an even more profound way to see it than the obvious differences that says something more powerful about the limits of computation than I've seen anyone say. Which is not to say somewhere out there someone already has had the idea I'm grasping for and written it in some brilliant paper or something. I'm just saying I haven't seen it.
"AI still hasn't passed a proper Turing test. A proper Turing test is 1 hour or more. Has any AI passed it?"
How do you answer that? "'Yes', AI has NOT passed a proper Turing test", or "'Yes', AI has passed a proper Turing test".
My comment doesn't actually make any prediction to judge, it's just an observation and an argument.
To your point, this example. The issue expressed here is with humans, not AI. We are still pretty terrible at writing specs. TBF, the AIs are too but that wasn’t being voted on.