Can the New Mathstral LLM Accurately Compare 9.11 and 9.9?
secondstate.io
secondstate.io
> In decimal representation, a number with a higher digit in the tenths place (the second digit after the decimal point)
The tenths place is the first digit after the decimal place.
You meant to write `1` I assume (as it claims "For 9.9, the first digit after the decimal point is also 1.").
No… they can’t. That’s like saying a search engine can solve math problems — which it can, in a sense.
I suspect that the people repeatedly saying this simply lack the knowledge to know what really constitutes a ‘complex math problem’.
And of course any half-decent new model can answer this particular question correctly; the designers aren’t stupid or unaware of what the expectations and common traps are. The model itself probably will be able to talk about why testing on such comparisons would be interesting (because it ‘knows’ about how this being a recent meme).
I was still incredibly impressed.
"Compare the first digit after the decimal point of both numbers.
- For 9.11, the first digit after the decimal point is 1.
- For 9.9, the first digit after the decimal point is also 1."
> "The 7B mathstral model answers the math common sense question perfectly with the correct reasoning."
Answers perfectly, sure. But the word "reasoning" is anthropomorphism and promises a level of cognitive ability that LLMs do not possess.
I know when a car has a flat tire, even if I don’t know a tappet from a carburettor.
It's good for everybody to know enough about how LLMs work to know why one shouldn't use an LLM to do math. Once that's understood, the need for and purpose of tools like Code Interpreter become clear.
Version 9.11 is greater than 9.9
Decimal 9.9 is greater than 9.11
Though that may just be because I program a lot more than I do arithmetic nowadays.
The question of what constitutes reasoning is extremely difficult to answer or attempt to answer in a rigorous way. We struggle with clear bright lines on this even as concerns biological organisms with vastly more flexible goal-directed behaviors, and on the closely related concept of consciousness, we lack a consensus even for humans in the womb.
Speaking for myself I tend to focus on “useful” levels of planning, generality, and goal-seeking to sidestep some of the thornier philosophical issues.
Even there things are wildly controversial. There is a significant group (including some very serious and credentialed experts) who claim that the trajectory is clear: some version of AGI is not only possible with these architectures but so imminent as to demand drastic policy decisions.
There is another group (likewise including unimpeachably credentialed experts) who claim that there is no evidence for this extraordinary claim, and that attention decoders in no way show potential for this kind of generality.
My understanding of the math and mechanism, for whatever it’s worth, inclines me to agree with the latter group.
The real answer is: no one could possibly know at this point. My impression is that very few experts believe that transformers alone will lead to AGI (which surely requires ‘reliable reasoning’), even amongst those who believe we’ll all be intellectually replaceable within a decade.
I meant approximators.
But I don’t think any of the big closed source models are relying entirely on next token prediction anymore. They are using reinforcement learning to add new (more complex) objectives to the training. This might allow for better reasoning abilities within the same architecture.
Probably not, but in some, or most, cases may be. You can see it in schools on math exams. Not all pupil can do it, some lack the knowledge, but most simply cannot put things together. "WHY" is another big question.
Wrong. GPT-4o gives me the correct answer to this question, 9.8.
(Note that the logic in the response from the LLM is blatantly nonsense).