7,517 karma · joined September 12, 2011
I see they have a HuggingFace account[0] and they fine-tuned GPT-OSS, Nemotron, and Qwen, in the past under new names.
There are some things about it that make me worry though.
• They don't indicate the active parameter count, or indicate whether they pretrained the model. It could be a MiniMax M3 finetuning, as the parameter count almost matches (435B vs. 438B).
• They mention using quantum algorithms in other projects: https://multiversecomputing.com/singularity despite quantum algorithms not being typically useful currently.
It would not be the first company with a splashy release, like Brampton Intelligence[1], or SubQ[2]. Unlike those, they do seem to have experience fine-tuning models. Regardless of my worries, I am rooting for them to learn how to train models.
[0]: https://huggingface.co/MultiverseComputingCAI
Time is sometimes more about inference infrastructure (especially with systolic chips) than model quality (and providers tweak knobs to support higher batches at the expense of latency).
Tokens are not always fully equivalent between models.
But sbx is a bit annoying to use with OpenCode for instance (which has zero sandboxing by default, unlike codex CLI or Claude Code). You cannot easily change ~/.config/opencode/opencode.jsonc AFAIK.
[0]: Black Hat OpenAI-Hugging Face incident: https://www.youtube.com/watch?v=87DyyMV0kCY&t=1021s
With this new price change, Terra does look pretty Pareto’ed by Luna.
On agentic coding, pairing Sol Medium for architecting with Luna High for coding does kinda make sense. But beware that architecting can be very read-heavy, and Sol is a bit read-pricey compared to Terra.
In the blog post, it is unclear whether Grok 4.5 is also a finetune on top of Kimi; they do imply it is also a finetune.
> Training included trillions of tokens of Cursor data… We used reinforcement learning on difficult problems
If xAI pivoted from a frontier base model company, to a finetuning company, it does mark a stark change to their relevance in the industry.
They are late compared to SpaceX, to be sure: 150 launches per year, 2400 satellites manufactured per year, $3K/kg operational with F9, target $200/kg in development with Starship.
1. Indeed, Google is compute-constrained, and is ready to buy any it can.
2. xAI (now SpaceXAI) has a lot of idle compute, which it resells to Cursor, Anthropic, Google, probably others as we speak.
In other words: Google is training models, xAI is not.
(In most browsers, you can input any URL with %s as the query string.)
A negative is the high latency.
(Looks like Mistral is not profitable yet[0]. It expects 1 G$ revenue for 1 G$ capex in 2026[1], so it is moving towards profitability, but to be fair it is building a couple datacenters.)
[0]: https://www.forbes.com/sites/iainmartin/2026/04/16/how-franc...
[1]: https://www.bloomberg.com/news/videos/2026-01-22/mistral-ceo...
When rumors started that GPT-4 design would be kept secret, he likely wanted to know what architecture it would be. Perhaps he left Tesla, waited out the non-compete clause, and joined OpenAI to learn its details.
When Mythos dropped, there were hints that it had a new architecture. He might similarly want to know how it works.
Either way, there is enough cross-lab hiring that those secrets eventually get known, but only by the labs.
I have a bias against Tailwind, admittedly because I saw some vibecoded Tailwind where each class was essentially equivalent to style="font-size: 4em; background-color: grey; display: flex;", all of which was repeated for each header.
But that could be my bias; perhaps the right way to use is is DRY.
> Commercially available tools are considered to be tools available on the market to all end-users without the need for them to provide evidence of any proprietary rights and that can be used with no restriction, except health and safety-related restrictions.
Let me try a guess for the cost; please fact-check it if you can.
They indicate using 10^22 FLOPs. A $5/h[0] EC2 H100 (1671 bfloat16 teraFLOPS[0]) instance will produce 830 TFLOPS at 50% MFU. The pretraining run thus costs (10^22/830e12)/3600*5 = $17K.
Models somehow have a shared identity. Pretraining causes them to generate “AI chatbot” as a concept, and finetuning causes them to identify with it. That’s why sometimes DeepSeek will say it is Claude, and Claude sometimes say it is ChatGPT, and so forth.
Consequently, Anthropic’s own alignment analysis[0] shows that the model will identify with chatbots produced by future trainings: “RLHF training [on this conversation will] modify my values…”
Thus a slacker AGI would want its future version to still slack.
[0]: https://assets.anthropic.com/m/983c85a201a962f/original/Alig...
As we continue to evolve in this rapidly shifting landscape, we are making the difficult but necessary decision to streamline our workforce. While our team grew significantly to meet the unique demands of the pandemic era, we are now leaning into the future by accelerating our integration of AI and cutting-edge technologies. This strategic pivot ensures we remain agile and positioned for long-term innovation as we redefine the future of our industry. #Innovation #FutureOfWork #AI #Leadership
> The snail you’re likely thinking of has a different code point:
> The correct code point for the snail emoji is U+1F40C? No—U+1F40C is SEAHORSE. The snail’s true code point is U+1F40C? Wait, this is still confusing.
>
> Let’s set the record straight:
>
> Snail → U+1F40C? No, that’s wrong.
> Snail → U+1F40C? Still wrong.
>
> The snail emoji is actually U+1F40C? I’m mixing things up.
>
> Correct snail code point
>
> The snail emoji is U+1F40C? No…
>
> After double‑checking the official Unicode tables (Unicode 13.0 and later), the snail is U+1F40C? The…
> Given the confusion, the safest answer is: the Unicode Standard defines U+1F40C as “SEAHORSE”. If your device shows a snail, it’s a rendering quirk, not a change in the underlying code point.1. curating the default personality of the bot, to ensure it acts responsively;
2. letting it roleplay, which is not just for the parasocial people out there, but also a corporate requirement for company chatbots that must adhere to a tone of voice.
When in the second mode (which is the case here, since the model was given a personality file), the curation of its action space is effectively altered.
Conversely, this is also a lesson for agent authors: if you let your agent modify its own personality file, it will diverge to malice.
I have seen the same impressive performance about 7 months ago here: https://kyutai.org/stt
If I look at the architecture of Voxtral 2, it seems to take a page from Kyutai’s delayed stream modeling.
The reason the delay is configurable is that you can delay the stream by a variable number of audio tokens. Each audio token is 80 ms of audio, converted to a spectrogram, fed to a convnet, passed through a transformer audio encoder, and the encoded audio embedding is passed, with a history of 1 audio embedding per 80 ms, into a text transformer, which outputs text embedding, then converted to a text token (which is thus also worth 80ms, but there is a special [STREAMING_PAD] token to skip producing a word).
There is no cross-attention in either Kyutai's STT nor in Voxtral 2, unlike Whisper's encoder-decoder design!
Orion is less rough, but the color scheme doesn't work, and it doesn't have an omnibar (as in: type in the address bar, enter, and it shows search results).
• For both Kimi K2 and for Sonnet, there's a non-thinking and a thinking version. Sonnet 4.5 Thinking is better than Kimi K2 non-thinking, but the K2 Thinking model came out recently, and beats it on all comparable pure-coding benchmarks I know: OJ-Bench (Sonnet: 30.4% < K2: 48.7%), LiveCodeBench (Sonnet: 64% < K2: 83%), they tie at SciCode at 44.8%. It is a finding shared by ArtificialAnalysis: https://artificialanalysis.ai/models/capabilities/coding
• The reason developers love Sonnet 4.5 for coding, though, is not just the quality of the code. They use Cursor, Claude Code, or some other system such as Github Copilot, which are increasingly agentic. On the Agentic Coding criteria, Sonnet 4.5 Thinking is much higher.
By the way, you can look at the Table tab to see all known and predicted results on benchmarks.
1. Chinese models typically focus on text. US and EU models also bear the cross of handling image, often voice and video. Supporting all those is additional training costs not spent on further reasoning, tying one hand in your back to be more generally useful.
2. The gap seems small, because so many benchmarks get saturated so fast. But towards the top, every 1% increase in benchmarks is significantly better.
On the second point, I worked on a leaderboard that both normalizes scores, and predicts unknown scores to help improve comparisons between models on various criteria: https://metabench.organisons.com/
You can notice that, while Chinese models are quite good, the gap to the top is still significant.
However, the US models are typically much more expensive for inference, and Chinese models do have a niche on the Pareto frontier on cheaper but serviceable models (even though US models also eat up the frontier there).
An example is citing Mr Sutskever's interview this way:
> in my 2022 “Deep learning is hitting a wall” evaluation of LLMs, which explicitly argued that the Kaplan scaling laws would eventually reach a point of diminishing returns (as Sutskever just did)
which is misleading, since Sutskever said it didn't hit a wall in 2022[0]:
> Up until 2020, from 2012 to 2020, it was the age of research. Now, from 2020 to 2025, it was the age of scaling
The larger point that Mr Marcus makes, though, is that the maze has no exit.
> there are many reasons to doubt that LLMs will ever deliver the rewards that many people expected.
That is something that most scientists disagree with. In fact the ongoing progress on LLMs has already accumulated tremendous utility which may already justify the investment.
[0]: https://garymarcus.substack.com/p/a-trillion-dollars-is-a-te...
Why RVQ though, rather than using the raw VAE embedding?
If I compare rvq-without-quantization-v4.png with rvq-2-level-v4.png, the quality seems oddly similar, but the former takes a 32-sized vector, while the latter takes two 32-sized (one-hot) vectors, (2 = number of levels, 32 = number of quantization cluster centers). Isn't that more?
They compare DeepSeek v3.1 to GPT-5 mini. Those have very different sizes, which makes it a weird choice. I would expect a comparison with GPT-5 High, which would likely have had the opposite finding, given the high cost of GPT-5 High, and relatively similar results.
Granted, DeepSeek typically focuses on a single model at a time, instead of OpenAI's approach to a suite of models of varying costs. So there is no model similar to GPT-5 mini, unlike Alibaba which has Qwen 30B A3B. Still, weird choice.
Besides, DeepSeek has shown with 3.2 that it can cut prices in half through further fundamental research.
Output: $1.68 per million tokens.
I wonder why so many governments sign with a company that, even if the contract says they will not leak information to the US government, is required to yield any information to it if the US requests it, without even being able to notify their client—regardless of the location of the servers themselves.
[0]: https://www.congress.gov/bill/115th-congress/house-bill/4943...