Not everything needs to be entertaining to be useful.
Not everything needs to be entertaining to be useful.
LLMs are bad at counting no matter what size of context is provided. If you're going to formulate a thought experiment to illustrate how an LLM stops paying attention well before the context limit, it should be an example that LLMs are known to be good at in smaller context sizes. Otherwise you might be entertaining but you're also misleading.
My hope from this article is to help non-AI experts figure out when they need to design around a flaw versus believe what's marketed.
You're putting a lot of weight into counting. I don't know anyone who wants to use a LLM after hearing "good at math" for counting of all things. Algebra, Calculus, Statistics, hell I used Claude 3 for Special Relativity. Those are the things people will care about when you say math, not counting.
Look, just test your use case and report that lol.
Your article would have been very helpful if you'd simply did that but you didn't so it's not.
https://www.lesswrong.com/posts/qy5dF7bQcFjSKaW58/bad-at-ari...