See The Open Source AI Definition from OSI: https://opensource.org/ai
See The Open Source AI Definition from OSI: https://opensource.org/ai
With LLMs, the list doesn't even have to be kept up to date, nor the links alive (though publishing content hashes would go a long way here). It's not like you can get an identical copy of a model built anyway, there's too much randomness at every stage in the process. But, as long as the details of cleanup and training are also open, a list of training material used would suffice - people would fetch parts of it, substitute other parts with equivalents that are open/unlicensed/available, add new sources of their own, and the resulting model should have similar characteristics to the OG one who we could, now, call "open source".
It's just not literally labelled so because of obvious reasons.
https://www.reddit.com/r/LocalLLaMA/comments/1ilsfb1/comment...
I recommend reading the actual Open Source AI Definition[1] and the FAQ[2]. There's also the whitepaper[3] that goes into much more detail about the state of affairs.
[1]: https://opensource.org/ai/open-source-ai-definition
[2]: https://hackmd.io/@opensourceinitiative/osaid-faq#What-is-th...
[3]: https://opensource.org/wp-content/uploads/2025/02/2025-OSI-D...
the model reveals the architecture which is all you need to use/run/train it.
For an LLM you can finetune and enhance, distill and embed given just the model weights, the runtime, and a permissive license. Having more is better. Well written detailed model release papers help a lot. Training code and training data are a great bonus.
However, I find the purity contest a bit too dismissive of the great contributions to the AI dev ecosystem that Meta and Deepseek have brought us. Without these, there wouldn't be the open ecosystem we have today.
If I write a program, then obfuscate it and then release the obfuscated code under an open source license, would you consider it open source(I would)? That's kind of the case here, they are releasing the model weights under an open source license.
Personally, I think it's fine to shorten it to "open source model" instead of "a model with the weights released under an open source license". What I would object to is releasing model weights under a restrictive license and calling that open source.
I wouldn't. Most definitions of open source say something like "in the form used for editing". You can release a built binary under an unrestrictive license, but that does not mean that you've opened the source. It's literally the plain meaning of the words: the source, as in where the thing comes from, needs to be open for it be meaningful.
But that's also true for binaries, games are a good example of where people pushed this quite far. Based on what little experience I have in ML, I'd say it's about the same thing. Whereas an API is more akin to a piece of software you can't tinker with in any way.
Guess the bar is just lower in the LLM space :P
Much in the same way, no sane company will touch the legal nightmare of releasing LLM training data scraped from public websites. Even releasing the LLM alone might be infringement, there are literally court cases being fought over this right now.
And that's fine! It's still valuable to have access to the source code, even if the "batteries" aren't included. Of course, if you really want to call it an open source model you should include the source for the data scraping/cleaning stages too; then the only thing missing would be the compute time and risk of acquiring dubiously-legal inputs.
I personally prefer a taxonomy like:
* Open weights: you can download the artifact and run it locally, not just use it through an application like chatgpt or an API.
* Open source: the code that created the artifact is provided in the same format that the authors used to work on it.
* Open data: the dataset that the source code was used on is available for download.
All three of those could be individually licensed or released, for 8 possible combinations. In the analogy to games, they would correspond to the licenses on the retail binary, the source code of the game, and the original uncompressed art assets or Blender projects, respectively.
I agree that open source doesn't mean open assets, but neither does open assets mean open source. You could make a linguistic argument that the training data is part of the "source" of the model (as in, from whence it came), but in any case the point is moot because neither the training data nor the code is open.
The AI crowd doesn’t care much for licenses anyway.