> tool call pruning breaks cache and people will tell you this is horrible and expensive
> except i looked at some anthropic data and real user behavior ends up with better cache hits and 30% less spend
> even this is needs to be analyzed further, it's just not simple
> for openai data it's inverted! cache hit ratio is actually better [sic: I think he meant worse based on the screenshot] with tool call pruning turned on
> but the net $ saved is only 5%
> kimi is a funny one - it has better cache hits with pruning on...but is also more expensive!
There was also another thread recently where he discussed that pruning improves user experience (models are smarter with less context) but I can't find it.
This can also be disabled in the config: https://opencode.ai/docs/config/#compaction
> our implementation is it only prunes calls from > 3 user messages ago, if context is > 40K, and only if there's at least 20K tokens to be removed
Seems reasonable to me and explains why I can have long sessions (way longer than with zed agents) while still hitting cache. Opencode is just missing per-provider TTL.
Ah, reminds me of good old "There are only 2 hard problems in computer science: cache invalidation, naming things, and off-by-1 errors."
You quip, but LLM KV caching (from the harness side) is quite easy: You get a cache hit on stable prompt prefixes, period. That means you want to keep the prefix stable, and only append at the end of the conversation. Made up example: Don't put the git branch name into the system prompt part (that comes first), as whenever the branch name changes, that'd trigger a cache invalidation of the entire prompt.
Getting this right requires some care to not by accident modify the prefix, basically, and some design on communicating the things that can change (user configuration, working dir, git information, ...).
Conceptually the underlying general idea is to sort things based on stability if you can avoid recomputing properties of the stable part.
They're aware of the issues length and they're "looking into a solution".
I want to know if I'm missing something cool!
At the end, cache hit rate is like 99.5% if Novita is not having issues.
For official DeepSeek API, 99.9% or something.
Custom harness that never compacts or otherwise doctors the history.
Usually I don't tell it to implement something adhoc, I first implement it in the documents first. LLMs are quite good to keep those documents in sync.
A good part of the implementation plan is that it keeps the LLM on track. With it, the LLM can understand why something must not be done yet, so it includes less unsolicited functionality. My workflow surely can be improved, but it has worked well for me.
In not sure about the actual costs, because I started using the same subscription for document parsing. But even then, I used less than $10 in may.
On the sheer performance it’s comparable to Opus ?
./cost.py amount-2026-5.csv 0.3 3.75 15
input_cache_hit_tokens: 472,971,520 tokens -> $141.8915
input_cache_miss_tokens: 13,299,013 tokens -> $49.8713
output_tokens: 3,334,962 tokens -> $50.0244
cache hit rate: 97.27% (472,971,520/486,270,533)
cache miss rate: 2.73% (13,299,013/486,270,533)
total: $241.7872
All of this usage was with an OpenCode subagent exclusively.Total input token = input + cache read + cache write Cache hit rate = cache read / total input token.
That is 71% in my very limited use of opencode.
The default is just 8GB and a full 128k context for the dense model can take most of that. So then comes an agent and causes eviction and subsequent cache miss.
Bumped the cache size (--cram IIRC) up to 48GB and had much better results.
I switched to vLLM and those went away. Need to look at my opencode config and adjust some others based on things I see here