Microsoft adding Deepseek support already as I recall?
That is - for any definition of "they are behind X months" then eventually they get to the point Claude was in January when the world freaked out, but at 1/10th the cost. A lot of firms are going to mandate that is good enough for their developers.
I believe this hasn't been confirmed yet but I think it speaks to a bigger problem for the AI companies which is, if you give capable developers a good reasoning LLM, they can make it work like it was a really expensive model.
I believe we are 100% at the stage of good enough for the vast majority of tech companines. Fable and others will be more valuable for non-traditional tech companies.
I read somewhere that the chinese AI companies are sharing knowledge and it would not surprise me if the government is applying pressure by saying work together or else. If they work together, they can truly commoditize LLMs and with China ramping up hardware support for AI, I see the future being inference speed and hardware being the moat.
Which makes sense to me. Selling a chatbot interface/model access to the general public was never going to be a viable long term play. You still need developers to wrap the models into specialized tools. Queue the Jobs quote "It's a feature, not a product."
I built my career on Solaris and it got rugpulled by Linux.
That wasn’t because of software, it was because of hardware. Linux’s cost advantage existed because Sun hardware had huge margins, because their software was basically free.
AI will probably be a repeat of this. Whoever can come up with the hardware solution that minimizes the cost per token will win.
I believe the 5090 still holds this crown, but someone certainly knows better than I do.
The only hiccup in that happening is will the US Gov let Anthropic and/or OpenAI fail when that time comes.
And of course the C-suite will have unlimited access to Mythos tier models, which they'll use to summarize reports, while passing down mandates to rank and file to increase usage of less expensive models.
OpenRouter charges an extra 5.5%, Fireworks does not, Google is separate, but I doubt it will take 18 months. They are already aware they are losing business.
OpenRouter is the wrong abstraction for enterprise, we only need one model provider, not everyone in the world. Nor do we want to have to worry about failover going to providers we don't want.
But that's because I never got on the "run three dozen agents in a ralph loop" trend or other high-token usage methods. The way I use AI is discrete and targeted and it seems that's how it will be for everyone once the economics settle.
At least on a personal I feel like I’ve been getting the same amount of work done but I have to think harder rather than sitting back and prompting and waiting.