Thinking Machines might be it.
Thinking Machines might be it.
Here are some of their current open weight offerings: https://www.arcee.ai/open-source-catalog
That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it better if it's just QLora or Lora) but let's see!
[0] https://thinkingmachines.ai/tinker/
[1] https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
What's their moat / secret sauce?
The Chinese "Neijuan" aside, most competing labs are going for the classic, 'your margin is my opportunity': https://tomtunguz.com/is-your-margin-my-opportunity-software... / https://archive.vn/5Vmq3
1. Magic
2. Managed hosting of their model
3. Hurting competitors. If people are using Meta’s commoditized models they’re not paying Google or allowing OpenAI to become too big.
4. Free R&D from open source. If open source developers are optimizing systems to run Llama, that helps Meta.
5. More magic
Well, he lost his job on that bet... and yet... I do not think that the verdict of history is quite in.
Feature as a company for now. Apple is struggling to build an in-house model set. And plenty of software behemoths, e.g. IBM, are realising they don't have a ticket to the new tech economy.
GLM-5.2 is the best in that class right now. It is competitive with current GPT/Claude/Gemini.
Benchmarks have GLM 5.2 somewhere underneath Sol and Fable and closer to now last-gen openai and anthropic models.
There's 2 caveats with the rest. First, GLM 5.2 matches those models in "xhigh" effort modes, which has a very low quota on the subscriptions, especially for Claude.
Second, last-gen GPT/Claude means what they release in April/May of 2026. Or to be even more complete/fair:
GLM 5.2 beats what OpenAI released in March 2026 (GPT 5.5 xxhigh), and what Anthropic released in April 2026 (Opus 4.7 xhigh). It is beaten by what OpenAI released in April of 2026 (GPT 5.6 Sol xxhigh) and Anthropic released in May 2026 (Opus 4.8 (the same as "Fable" ?), xhigh effort)
GLM 5.2 was released on Jun 16 and if OpenAI and Anthropic hadn't done those quick releases they would have been beaten on their best available models ...
So great news! Open source now has SOTA performance 3 months after OpenAI/Anthropic/Google. Wow.
Opus 4.8 (May) to Kimi K3 (July) has apparently just dropped it to two months.
China also does efficiency improvements. Qwen 3.6 27B is better than Sonnet 4.5 and you can run it on a couple of gaming video cards. That's incredible. I can do real actual work with this!
As Google said in 2023, none of them have a moat, open weight models will win.
Google, who probably canceled the release of Gemini 3.5 pro to avoid having their best new model perform worse than BOTH GLM 5.2 AND Kimi K3?
I mean how is this anything but an incredible defeat of Google's supposed inventiveness?
Newly released Kimi K3 is benching better than Claude Opus 4.8. The only better models are Claude Fable and GPT 5.6 Sol Max Effort.
my bet is that Chinese government fund Chinese models way more compared to what those companies receive (except llama, which is outdated but was strong foundation at its time)
I think the bet would have to be that a US Open Weight company either: 1. Gets a lot of money from Jenson who views them as a counterbalance to the big labs in his ecosystem and a way to generate leverage (the same way he is positioning neoclouds-- it also could be synergistic with neoclouds who could offer the model serving endpoints) 2. Can fast follow the same way Mistral does (which, honestly, seems like just distilling the Chinese model, which distills the US lab but is pretty innovative on a whole lot of architecture both in training and serving land.) 3. AND figure out some (maybe not super lucrative but lucrative enough) sort of business model, as well. There are lots of possible business models, so I will be curious how this whole space evolves.
I suspect 2B is not enough to boostrap frontier model from the scratch (for both talent and hardware)
You can pretty much remove the supposedly here
we will see!
I find it wonderful that, as a non-profit, they are only one to two years behind SOTA models that cost billions of dollars to build, if not more.
Hopefully it somehow works out though!
Open-source models + services. This is more attractive because it doesn't lock in the vendors. If I grow larger, I can decide to deploy the open-source models.
Tech history is littered with the corpses of "open source but we sell hosting" services. Models are so expensive to train, you can't be losing the big clients once they get super profitable.
I get that they're in very different businesses, but for both don't they have the issue that once a client gets big enough the client might decide to move the services in-house? Based on how much of the internet went down when that AWS data center crashed the answer is clearly "No" for AWS.
Is that because of physical, real-world infrastructure? Are there no open versions of their APIs? Is it too hard to migrate to something else once a client has achieved that size?
I would say "it's risky and requires a lot of labor to migrate without corruption, loss of data" and also minimizing downtime. Sure anyone can run pg_backup, but can you do it across 90 databases? Can you do it live? Can you coordinate rollout of the process, cutover, and monitor for failure? What's the cost of egress for this? Is the team your A-team or the B-team? Can you trust this to the B-team? Is it worth having this team spend all this time on a migration rather than, say, getting something new set up, or optimizing performance on an existing system?
I'm a database guy, but the same migration argument is presumably also extra work for (say) blob storage, networking, etc.
Since LLMs are stateless by their current implementation, switching to "the same open-weight model running in a different datacenter run by a different vendor" is "just" switching the API endpoint. (If they are the exact same shape, it's fine, if they differ somehow, there's perhaps some work to do there, fixing things and monitoring for failures on switch-over)
There are several open APIs it seems and OpenRouter.ai is doing a fine job making a commodity out of models and datacenters.
Database is more difficult, but tons of people have done it successfully.... meanwhile people who host their own LLMs are relatively small in number in comparison.
Most companies don't do their own data centers mainly because it is more expensive and less reliable. It's something they can just pay for the problem to go away. The calculus for hosting your own LLM is probably similar.
Even Stripe who built their own coding agents and has tons of money/resources still decides not to host their own LLMs.
Still, many people will prefer open-weight models. It is similar to how we prefer linux but still use AWS/Render/and whatever. It doesn't lock us in, and we can move providers if we want to.
AWS owns the hardware, and doesn't write a lot of the software.
AWS actually is kind of the opposite - it often takes open source software (e.g. Apache, Mongo, Kubernetes) and then makes money off it by hosting it itself (with some enhancements etc).
If they do develop their own software (e.g. with S3) they don't give away the source code so others can deploy it, as that's part of their secret sauce.
In this scenario, where they would be offering the open source model and then offering the same model hosted, there isn't really a moat here - they would be leasing the hardware from a company like AWS, and adding a margin, but it woudl be trivial for another company (or Amazon) to take their same model and offer it for the same price or less.
there is a chance their business model is absorbing government funding..
FWIW this is the same logic for China’s need for their own OW models
If you understand the world through a Chinese LLM, you are seeing it through a biased lens stemming from biased training data.
(Also, in that way, having all major LLMs developed by the US carries a risk too. We need more diversity than just the viewpoints of the US or China.)
Frankly the EU and the US will practically be less involved and have more pushback from the public in this than China. I think that’s less “China bad” than recognizing that China is a more authoritarian state and has far more proclivity to interfere than western states.
Maybe I’m wrong? What does deep seek say about Tiananmen square in 1989?
This is what Deepseek replied when I asked it with a burner account. Claims it doesn't have it in its training data... sure.
Every model will have their own bias. Freedom is relative and American freedom is not the only form of freedom. China bad or China authoritarian and thus not free is a result of the red scare
"Surveys of genocide scholars (e.g., one by the journal Journal of Genocide Research contributors) show meaningful disagreement, though a visible and growing share of specialists in genocide studies specifically have concluded the term applies or is defensible — more so than international law scholars generally, who tend to be more cautious about the intent requirement."
You can be of the opinion that this shows bias, but it's a far cry from the Chinese censorship.
Clichés like "freedom is relative" are not serious arguments. You can't genuinely argue that the US, for all its problems, is less free than China. Words have meaning.
E.g. looking at freedom house index you can disagree on the precise score or how different aspects are weighed, but you can't argue with the underlying facts in the country reports
You can't genuinely argue that the US, for all its problems, is less free than China
This is how Americans convince themselves they are at the top of the world when they are not. No one outside of Europe and the US cares about the US anymore. They buy Chinese products and do business with Chinese businesses instead, right now. Worldwide, China is seen more positively than the US specially in global south countries. The exceptions of course are Europe and The US, the cold war first world countries of course.
freedom house index
This is obvious American bias. Astroturfing.
There is also AllenAi in the US, but they have yet to produce a model at this scale. Thankfully, new contenders can come out of nowhere and do well, as long as they can produce a competitive model.
GLM 5.2 underwent extensive post-training and iteration since its original release to reach its current state. This seems like an extremely strong model for a first release, with a lot of potential for improvement, just like DS4.
Sometimes I wish Meta had stuck with Llama 4 a bit longer to see how much further it could be pushed.
They overspent on llama 3 anyway so money ran dry, LeCun is good at running research, but budgets didn't stretch. Meta isn't investing in frontier big models anymore.
Yes they are. Meta Muse is their attempt.
It's below frontier performance at the moment but they are spending on getting there.
Meta Spark is moderately promising but of course closed source.
What does "winning" mean to you?
There’s also Prism