We need:
1. Open datasets for pretrains, including the tooling used to label and maintain
2. Open model, training, and inference code. Ideally with the research paper that guides the understanding of the approach and results. (Typically we have the latter, but I've seen some cases where that's omitted.)
3. Open pretrained foundation model weights, fine tunes, etc.
Open AI = Data + Code + Paper + Weights
These datasets are huge, and it's practically impossible to make sure they are clean of illegal or embarrassing stuff.
As far as I know the only true open source model that is competitive is the OLMo 2 model from AI2:
https://allenai.org/blog/olmo2
They even released an app recently, which is also open source, that does on-device inference:
https://allenai.org/blog/olmoe-app
They also have this other model called Tülu 3, which outperforms DeepSeek V3:
Lets say you took GCC, modified its sources, compiled your code with it and released your binaries along with modified GCC source code. And you are claiming that your software is open source. Well, it wouldn’t be.
Releasing training data is extremely hard, as licensing and redistribution rights for that data are difficult to tackle. And it is not clear, what exactly are the benefits in releasing it.
It's a return to the FREEWARE / SHAREWARE model.
This is the language we need to use for "open" weights.