How is this comparable to going to lunch or taking a walk?
If you keep this running for hours without doing anything, it will drain your limits and API. The use case of keeping the main thread cache warm while subagents work is very genuine and legitimate.
Keeping your cache warm is a good thing, caching saves compute and electricity.
Cached input is cheap for a reason, it is in everyone’s mutual interests to maximise cache hit rates.
The cache is discounted for a reason. They WANT you to use it.
1) Start charging for VRAM reservations.
2) Charge _other_ customers more.
3) Eat the cost themselves.
Anthropic (and now OpenAI too for 5.6) prompt caching is not free.
"They already charge me to park my car, why can't I leave it there for a year for the same price as 1 week?"
I don’t think your understanding works.