> But without more information about the training data and methodology, it isn't exactly "open source".
Being or not being open source has exactly jack shit to do with that.
Being or not being open source has exactly jack shit to do with that.
The idea behind OSS is that you're able to modify it yourself and then use it again from that point. With software, we enable this by making the source code public, and include instructions for how to build/run the project. Then I can achieve this.
But with these "OSS" models, I cannot do this. I don't have the training data and I don't have the training workflow/setup they used for training the model. All they give me is the model itself.
Similar to how "You can't see the source but here is a binary" wouldn't be called OSS, it feels slightly unfair to call LLM models being distributed this way OSS.