Well technically even DeepSeek is not as OSS as OLMo or Open Euro, because they didn't open the data.
We need:
1. Open datasets for pretrains, including the tooling used to label and maintain
2. Open model, training, and inference code. Ideally with the research paper that guides the understanding of the approach and results. (Typically we have the latter, but I've seen some cases where that's omitted.)
3. Open pretrained foundation model weights, fine tunes, etc.
Open AI = Data + Code + Paper + Weights
These datasets are huge, and it's practically impossible to make sure they are clean of illegal or embarrassing stuff.