See the response here: https://pastecode.io/s/6uyjstrt
See the response here: https://pastecode.io/s/6uyjstrt
0 - 1
5 - 0
6 - 1
7 - 0
8 - 2
9 - ?
Grownups try to find a pattern in the numbers, different types of series, progressions, etc. The correct answer is 1 because it's the number of circles in the graphical image of the number "9".Token 145 (ar) = Token 236 (a) + Token 976 (r)
Repeat many times with different combinations and different words?
Sure, but then we would lose a benchmark to measure progress of emergent behavior.
The goal is not to add one capability at a time by hand - because this doesn’t scale and we would never finish. The goal is that it picks up new capabilities automatically, all on its own.
For example, why wouldn't "0" be a correct answer here (rule being "every other number on the right should be 0, other numbers do not have a pattern")?
In a sibling comment, I replied that usually a repeating pattern can also be applied, but that one usually requires storing the n-sequence, rarely making it the shortest encodable rule.
A rule being "good" is largely about simplicity, which is also essentially the trick that deep learning uses to escape no-free-lunch theorems.
Or just, the pattern simply repeats for the n next number.. like come on, there is no objective metric for which rule is "better".
An extension of the idea could be that many other similar tests trying to measure and/or evaluate machine cognition, when the LLMs fails, are not precisely measuring and/or evaluating anything else than an specific edge case in which machine cognitions fails (i.e. for the specific LLM / AI system being evaluated).
Maybe the models are actually more intelligent than they seem, like an adult failing the number of circles inside the graphical image of the numbers, in the mentioned problem.
[...]
Starting with 0 paired with 1. If I think about functions or some formula that could map these values, but that might be too complicated.
Alternatively, perhaps these numbers represent something specific. For example, maybe they're representing the number of certain features in the number itself, like the number of closed loops in the digit's shape.
Let's think about that:
0 has one closed loop. 5 has no closed loops. 6 has one closed loop. 7 has no closed loops. 8 has two closed loops. 9 has one closed loop. Wait a minute, that seems to match the sequence:
0 (1 loop) - 1
5 (0 loops) - 0
6 (1 loop) - 1
7 (0 loops) - 0
8 (2 loops) - 2
So, if 9 has one closed loop, then it should be paired with 1.
But let's double-check this theory with the given pairs.
0 has one loop, paired with 1 – matches.
5 has no loops, paired with 0 – matches.
6 has one loop, paired with 1 – matches.
7 has no loops, paired with 0 – matches.
8 has two loops, paired with 2 – matches.
Yes, this seems consistent.
Therefore, 9 should have one closed loop, so it should be paired with 1.
[...]
Then it continued trying another approach related to the number of angles in the digit shape, but it discarded it. Then "Alternatively, perhaps it's related to the number of segments that are 'on' in a seven-segment display." and so on.
thinking about the phonetics, "strawberry" has three /r/ sounds...
For example, in Spanish, it's "fresa," which has no "r"s, or in French, "fraise," which also has no "R"s...
But to be absolutely certain, let's count them one by one: s-t-r-a-w-b-e-r-r-y. Okay, the first "r" is after the "a," then after the "b," and finally twice at the end before the "y." Wait, is that two or three "r"s at the end? Let's see: s-t-r-a-w-b-e-r-r-y. So, r-a-w-b-e-r-r-y. Okay, there are two "r"s before the "y." So, that's one after "a," one after "b," and two before "y," making it four "r"s in total. Wait a minute, now I'm confusing myself...
For example, "copyright" has two "r"s, "arroba" has one "r,"...
Wait, perhaps I'm confusing it with "arrowroot," which has two "r"s. Alternatively, maybe thinking of a word like "corroborate," which has four "r"s...
and Alternatively, thinking about Middle English "strawberie," still three "r"s.
It's not being thorough. It's just trained to sound like it's being thorough. It's producing the same garbage as any other LLM, this one just happens to get the answer right because it's been trained on data that includes discussions of the strawberry problem.There's certainly a unique mechanism which might make it better at searching a space for answers, and that's pretty neat in its own way.
An LLM trying to get the number of letters will just be regurgitating for the most part because afaik it has no way to actually count letters. If the architecture was changed to allow for this (breaking certain words down into their letter tokens rather than whole word tokens) then it may help, but is it worth it?
The tone and expression is novel and it _looks_ like there’s something fundamentally different about reasoning but… also it keeps repeating the same things, sometimes in succession (a paragraph about “foreign languages” then another about “different languages”), most paragraphs have a theory then a rebuttal that doesn’t quite answer why the theory is irrelevant, and sometimes it’s flat out wrong (no Rs in “fraise” or “fresa”?).
So… holding my judgement on whether this model actually is useful in novel ways