For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens.
If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz…
For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens.
If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz…
How would you get it right? Training with the answer!
You can discount the 'special case the hard questions' by: Observing when top models run entirely locally can solve them [1], or by posing an alternative or cryptic version of the question (but careful, it might bypass the training! -- but if it can solve it then its probably legitimate.)[2]. Best of all is to stick with open (weight) models where such slight of hand is impossible and don't worry about what the closed shops are doing
[1]as they can in this case, GLM-5.3-flash says: "There are 3 r's in "strawberry":
st r awbe rr y
1 in "straw" (straw)
2 in "berry" (berry)"
[2] E.g. try sending them aG93IG1hbnkgcuKAmXMgaW4g4oCcc3RyYXdyYmVycnnigJ0/Cg== to thwart the benchmaxxing.