https://openai.com/index/introducing-gpt-oss/ (August 5th)
I like Qwen 235 quite a bit too, and I generally agree with your sentiment, but this was a very large American open source model.
Unless we're getting into the complications on what "open source" model actually means, in which case I have no clue if these are just open weight or what.
Inference is usually less gpu-compute heavy, but much more gpu-vram heavy pound-for-pound compared to training. General rule of thumb is that you need 20x more vram for training a model with X params, than for inference for that same size model. So assuming batch size b, then serving more than 20*b users would tilt vram use on the side of inference.
This isn't really accurate; it's an extremely rough rule of thumb and ignores a lot of stuff. But it's important to point out that inference is quickly adding to costs for all AI companies. Deepseek claims that they used $5.6mil to train Deepseek R1; that's about 10-20 trillion tokens at their current pricing- or 1 million users sending just 100 requests at full context size.
Not when you have to scale. There's a reason why every LLM SaaS aggressively rate limits and even then still experiences regular outages.
There is so much misinformation both on HN, and in this very thread about LLMs and GPUs and cloud and it's exhausting trying to call it out all the time - especially when it's happening from folks who are considered "respected" in the field.
they may be taking some western models: llama, chatgpt-oss, gemma, mistral, etc, and do postraining, which required way less resources.
The NYT used tricks like this as part of their lawsuit against OpenAI: page 30 onwards of https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
Maybe I'm wrong about that, but I've never heard any of the AI training experts (and they're a talkative bunch) raise that as a suspicion.
There have been allegations of distillation - where models are partially trained on output from other models, eg using OpenAI models to generate training data for DeepSeek. That's not the same as starting with open model weights and training on those - until recently (gpt-oss) OpenAI didn't release their model weights.
I don't think OpenAI ever released evidence that DeepSeek had distilled from their models, that story seemed to fizzle out. It got a mention in a congressional investigation though: https://cyberscoop.com/deepseek-house-ccp-committee-report-n...
> An unnamed OpenAI executive is quoted in a letter to the committee, claiming that an internal review found that “DeepSeek employees circumvented guardrails in OpenAI’s models to extract reasoning outputs, which can be used in a technique known as ‘distillation’ to accelerate the development of advanced model reasoning capabilities at a lower cost.”
there was obviously llama.
Shenzhen 2025 https://imgur.com/a/r6tBkN3
https://ifiwaspolitical.substack.com/p/euroai-europes-path-t...
This feels like a joke... Parity with a 2024 model in 2027? The Chinese didn't wait, they just did it.
The timeline for #1 LLM is also so far into the future that it is entirely plausible that by 2031, nobody uses transformer based LLMs as we know them today anymore. For reference: The attention paper is only 8 years old. Some wild new architecture could come out in that time that makes catching up meaningless.
GPT4 parity on a own silicon trained indigenous model is just an early goal.
Indeed, the ultimate goal is EU LLM supremacy - which means under democratic control.
Well, that's true... but also nobody else is. Making something popular isn't particularly impressive.
It's kind of like releasing a 3d scene rendered to a JPG vs actually providing someone with the assets.
You can still use it, and it's possible to fine-tune it, but it's not really the same. There's tremendous soft power in deciding LLM alignment and material emphasis. As these things become more incorporated into education, for instance, the ability to frame "we don't talk about ba sing se" issues are going to be tremendously powerful.
And how would releasing open-weight models help with that? Open-weights invite self-hosting, or worse, hosting by werstern GPUaaS companies.
It’s true that DeepSeek won’t give you reliable info on Tiananmen Square but I would argue that’s a very rare use case in practice. Most people will be writing boilerplate code or summarizing mundane emails.
Deepseek 3.2 is 1% the cost of Claude and 90% of the quality
>“We believe the benefits of superintelligence should be shared with the world as broadly as possible. That said, superintelligence will raise novel safety concerns. We’ll need to be rigorous about mitigating these risks and careful about what we choose to open source.” -Mark Zuckerberg
Meta has shown us daily that they have no interest in protecting anything but their profits. They certainly don't intend to protect people from the harm their technology may do.
They just know that saying "this is profitable enough for us to keep it proprietary and restrict it to our own paid ecosystem" will make the enthusiasts running local Llama models mad at them.
1) The four models you mentioned, combined
or
2) ChatGPT
?
What gives? Because if people are willing to pay you, you don't say "ok I don't want your money I'll provide my service for free."
Like research labs and so on. Even at US universities
Now you have the answer to "what gives" above.
Best they can hope for is getting acquired by MS for pennies when this scheme collapses.