You a are right as The article doesn't mention KV cache. Also, yes the writing is slop, and does not mention the LLM special tokens. Only discusses the use through the OpenAI-style JSON wrapper that allows you to define the schema of a tool call in JSON. For most of the audience, they are looking to build AI agents, and the high-level tool calling interface is what they would be using