Well LLMs are claimed to be good at math too, and yet they can't count. Same point with the long contexts. And our actual use case (insurance) does need it to do both.
My hope from this article is to help non-AI experts figure out when they need to design around a flaw versus believe what's marketed.