5 karma · joined May 7, 2026
But I guess for Agents I would continue using AgentCore.
Maybe if I wanted to build my own AgentCore though.
This way we can save Anthropic and co. for example the effort of recalculating all the linear algebra needed for the prefill of the system prompt, for which they reward us with reduced input token cost. The result is the same, if cached or not cached, it's just less computation.
For the same prompt you would still get the same answer (assuming temperature = 0).
The big savings you will get in an agentic/conversational context: Every new turn always puts the full message array back into the request. If we don't cache the calculation result at every step, the provider has to recalculate the early turns potentially hundreds of times (see second/third graphic).
I will never go back. Lets see if in 3 years I have a different opinion, but I can't imagine so.