General Reasoning: Free, open resource for building large reasoning models
gr.inc
gr.inc
- The initial use of data is distillation so we’re less bound by question quality (anything that evinces output diversity is good).
- But moving onto RL, we’ll need stronger quality. We have much better things planned both on data filtering and verification!
- Surprisingly, a lot of ML datasets actually look like this when you look under hood. We’re hoping having more eyeballs on it will help improve quality in long run over less transparent status quo!
It makes sense for ASI research I suppose, but why are we trying to teach small models to do stuff almost no humans even try to do?
What happens if you train them with RAG context in the prompts and calculator calls in the CoT?
I agree with your meta-point that better benchmarks testing more types of task would be good!
Can LLMs apply a consistent procedure for logic puzzles with logically disjunctive possibilities?
Enter: Philosoraptor the LLM
Once this stage is reached, once we can throw piles of data onto a reasoning model and get formal algorithms that explain or predict that data, the new era will begin.
The concepts are heavily covered in the training corpus, and if people were allowed to take it more than once, with even a book let alone access to the internet it wouldn't be very hard.
Examples:
1) Find the sum of all integer bases $b>9$ for which $17_b$ is a divisor of $97_b.$
In the corpus: https://www.quora.com/In-what-bases-b-does-b-7-divide-into-9...
And one more:
3) https://artofproblemsolving.com/wiki/index.php/2025_AIME_I_P...
Is just the the number of ways to distribute k indistinguishable balls (players) into n distinguishable boxes (flavors, without exclusion, in such a way that no box is empty.
Thus in the corpus for any courses that need to cover combinatorial problems including physics, discreet math, logistics etc...
IMHO these concept classes from a typical AIME are so common, the scores you gave demonstrate that those models are doing no "general reasoning" at all and are actually failing at approximate retrieval.
(Also the term “approximate retrieval” is a bad one - reasoning is inherently a process of chaining together associations. What matters is whether the reasoning reaches the right conclusions. Still some way to go, but already very impressive in tasks traditionally considered harbours of human reasoning!)
It would have been seen as witchcraft.
no, it doesn't. a broken clock is right twice a day, reasoning is about the journey more than the destination
good luck. as a reminder, there are people who, with varying degrees of certainty, think their loved ones have been replaced by actors, as well as people who think they're actually the god of the world around them, for it is just their imagination.
None of that is any proof at all that LLMs or computers in general can reason.
"some humans are dumb, so LLMs are smart" is not a valid argument here.
Not true.
A LLM might give you that answer x% of the time, x being a number less than 100. However, any thinking person answering your question, will give you the same answer, no matter how many times you ask it. That's the fundamental difference between thinking and statistically mapping and reproducing the structure of human language.
Oh, will they? Will they really?!
Never mind that the tendency to give the exact same answer to the same question over time is not the exhibition of reasoning power you seem to think it is. Have you actually asked some people to multiply 10-digit numbers in their heads? Did they always get the same result? No? Well, there goes that argument.
We don't do anything that the LLMs don't do at this point, except adjust our weights (poorly) to move short-term context into long-term memory. Once that capability is added to the models -- which will happen soon enough, because why wouldn't it? -- where will the goalposts go next?
Of course those who don't care about improving those systems also don't care about understanding their limits, which is unsurprisingly the case for a lot of people on this website.
Heck, let's throw in a square root for the fun of it: https://i.imgur.com/Q9eHAaI.png
How'd it do that, if it can't reason? That problem wasn't in its training corpus. Similar ones were, with different numbers, and that was enough.
Ask it 100 times, and it will probably get it wrong a certain percentage of the time... just like you would if I asked you to perform the calculation in your head.
Notice that the model actually got the LSD slightly wrong in this example. 188.58 would be a better estimate. It even screws up the way we do. That, to me, is almost as interesting as the fact that it can deal with the problem at all.
Of course those who don't care about improving those systems also don't care about understanding their limits
The people who do care about improving these systems seem to be doing a pretty awesome job.
As for the limits, they frankly don't seem to exist. They certainly aren't where you and your predecessors over the past few years have assured us they are.
Cause I have, and I have doubts it was significantly more complicated math than basic integer addition and subtraction. I also do use a calculator even for basic, low value, integer math, because what do you know, in my perfectness I often had an issue with numbers not ending up as what they were supposed to.
There's also the quite accessible and extensive history of human calculators and the extensive error correction strategies they had to employ because they'd keep cocking up calculations.
Come on...
It was never meant to be a proof that LLMs or computers in general can reason [or rather, that they can reason generalistically]. Instead, it was a demonstration of how their argument looks when mapped to other situations, illustrating that it isn't really putting anything to the table in the ways of, proofs, evidences, definitions, or logic / arguments, nor does it enable others to do so.
> some humans are dumb, so LLMs are smart
That kind of reading loses the point just about completely, so it shouldn't be surprising you can't find an argument in there.
The point was and is that simply pointing at an AI model and saying "nah it's cappin" is awfully lacking, and is suspiciously similar to how certain people with certain mental conditions view their world. It is not insightful, nor reasonable. It's just an assertion of disbelief, following a - as you can hopefully agree - dubious logic that cannot be disproven or substantially argued against, as it was never designed to enable that in the first place.