If you read DeepSeek's papers, you'll find a litany of architectural features that allow for a greatly reduced cache hit price by shrinking the size of the KV-cache.
They just aren't in any hurry to forward those cost savings to you.
It very much possible Anthropic, OpenAI and Google able to serve their models much cheaper than their current API prices.
They just dont do it because they dont try to ubdercut each other and so far chinese models been percieved behind SOTA.