https://martinalderson.com/posts/watch-out-for-cache-read-co...
Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases).
Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.
edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.
These cache Hit % are accurate, I've done a ton of testing of this myself. The cache hit % is one of the most important metrics as far as estimating cost. There are many providers with cheap cache reads, but have an effective cache hit % of 30%, making their cheaper cache pricing meaningless compared to another provider who charges more but has a 85% cache hit percentage.
[0]: https://openrouter.ai/deepseek/deepseek-v4-flash-0731?endpoi...
scroll down on the provider/model card and you'll see a field called cache hit %, its different for every provider/model.
I don't use routing on openrouter, I strictly use models with a single provider and no fallback, at least for use with harnesses its pretty dumb to route requests to multiple providers you are busting your cache every other request and increasing costs by 20-50%.
FWIW, I get significantly higher than listed cache hit rates when I pin my session to a specific provider, which is further evidence of the above.
This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.
I'm not sure it's wholey accurate to say they "randomize" the provider, rather my assumption based on usage is that it's something like cheapest-ish/responded to the request within some reasonable-ish time/etc algorithm that chooses the provider on each request - which seems, remarkably questionable in terms of optimizing for user experience or hidden user costs.
> This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.
I so very much recommend this approach. My avenues that automate llm calls to openrouter are setup to make api reqs to openrouter to determine best price/response/etc and then pin the request to that (and, preferably, a fallback if there's reasonable difference between #1 and #2) provider for that session. Otherwise you're going to have a bad time.
I'd imagine this could make things interesting in cases where one provider is offering different quants than the others and openrouter is just swapping you back and forth on a long agentic session.
I don't believe this is correct? AFAIK once it routes you to a provider for a given conversation that choice is sticky unless you hit technical difficulties. (It's more complicated than that, they recently added named routing strategies that you can append to the model name.)
IMO the relevant metric is cache TTL which isn't typically published AFAIK.
I don't know why there's no "Pick the cheapest provider above nTPS on first request and stick until cache bust" setting.
I thought you had to actively manage caches, do you not?
Is there some stochastic process that takes place during those 5 minutes that determines whether or not you get the discount?
is this true?
Caching was always here, you don't need to do anything special to get it on a single user local backend running a base model or a chatbot in the first place. Among commercial providers, OpenAI adopted it in 4o first.
There are two problems here:
- cache hit pricing (both Muse Spark 1.2 Contributor and MiMo 2.5 are around the $0.002-3/M mark)
- cache persistence time
Muse Spark drops the cache in less than 5m. MiMo keeps it around for at least an hour based on my experience with whoever is serving it for OpenCode. This difference itself will inflate bills massively.
A 500K token input repeatedly read by MS 1.2 for full input price 12 times an hour = $0.60. You would be expecting $0.012. So a 50x difference. Same thing on MiMo 2.5 is $0.018 because of longer cache times.
it is basically the old dsv4-flash prices, but even more smart.
MiMo wins handsomely if you want to think about your code for minutes at a time as you write. I use it to make changes as I think. I know it will screw up some stuff. I then switch to MS/DS4 once every few hours and have it do a code review and fix the broken stuff. So much cheaper than getting MS to do it on its own.
I guess it could be fake but seems more likely people are just trying it out. Hy3 was a very strong and underrated model.
Its already serving as much tokens/day as the incredibly cheap and good GLM 5.3 flash, which had a crazy marketing campaign as ox alpha?
Also those top 5 apps are just 1.58B tokens out of 1.54T tokens from yesterday. Negligible.
wouldn't trust they dont do Capitalism like the rest of the AI field.
Like lobbying the US president to harm their competitors?
Administration corrupted up to it's very core? Check.
Nihilism of anyone not part of the proper color, gender, whatever agenda? Check.
Unlawful surveillance? Check.
Sending totally innocent citizens to prison with many of them dying mysteriously? Check.
Killing innocent people in the streets simply because they dare protest peacefully? Check.
Welcome to North Korea!
Oups. Confused.
Welcome to the GREAT US of A! Where True Capitalism is practiced.
Better yet, just go back to Reddit.
Authoritarian capitalists exist. Larry Ellison, Alex Karp, Elon Musk, Mark Zuckerberg - they love when you think that capitalism plays by written rules. It helps them use your tax dollars to pay for their surveillance services. It helps them sell you B2C products that degrade the moral fabric of America, and rewrite legislation through online influence campaigns. You can say it "isn't capitalism" because you don't agree with it politically, but the accrual of capital is precisely how all 4 of those men became powerful. Their influence on social media is why they're able to reframe the national discussion how they want. Their money is what lets them turn political willpower into a business, or a rubber chicken into a politician.
After all, what is Reddit or 4chan if not one big capitalist influence campaign? Do you really think product reviews and political ragebait is driven by some natural and benevolent social force that we can't see or feel?