I hate to be that guy, but this is (a) little to do with the actual problem at hand in the article, and (b) a dramatic oversimplification of the real challenges with LLMs.
> LLMs don't think. At all. They do next token prediction.
This is very often repeated by non-experts as a way to dismiss the capabilities of LLMs as some kind of a mirage. It would be so convenient if it were true. You have to define what 'think' means; once you do, you will find it more difficult to make such a statement. If you consider 'think' to be developing an internal representation of the query, drawing connections to other related concepts, and then checking your own answer, then there is significant empirical evidence to support high-performing LLMs do the first two, and one can make a good argument that test-time inference does a half-adequate, albeit inefficient, version of the latter. Whether LLMs will achieve human-level efficiency with these three things is another question entirely.
> If they are conditioned on a large data set of people repeating the same seven knock knock jokes over and over and over in some complex pattern (e.g. every third time, in French), what they produced will look like that, and nothing like thinking.
Absolutely, but this has little to do with your claim. If you narrow the data distribution, the model cannot develop appropriate language embeddings to do much of anything. You could even prove this mathematically with high probability statements.
> Failing to recognize this is going to get someone killed, if it hasn't already.
The real problem as in the article is that the LLM failed to intuit context, or to ask a followup. While a doctor would never have made this mistake, the doctor would know the relevant context since the patient came to see them in the first place. If you had a generic knowledgeable human acting as a resource bank that was asked the same question AND requested to provide nothing irrelevant, I can see a similar response being made. To me, the bigger issue is that there are consequences to easy access to esoteric information for the general public, and this would be reflected more in how we perform reinforcement learning to assert LLM behavior.