How come no other big model seems to be able to deliver the same type of extremely low cache cost though, if their techniques are public?
It very much possible Anthropic, OpenAI and Google able to serve their models much cheaper than their current API prices.
They just dont do it because they dont try to ubdercut each other and so far chinese models been percieved behind SOTA.
They just aren't in any hurry to forward those cost savings to you.