Today's open source models beat closed source models from 1.5 years ago
old.reddit.com
old.reddit.com
Although training the actual model is not feasible for people without the funds and hardware, having those things would at least allow the model to be auditable. Otherwise, we really have no idea what these open weight models are doing. They could be biasing themselves in various ways that are invisible but damaging.
Another thing to watch out for is licensing. Only a few models use an actual open source license from OSI (like Apache). Many are using proprietary licenses that reference external terms that can change over time, or limit how you use their model. These restrictions are definitely not in the spirit of open source.
In other words, much of this is “openwashing”, which is a new trend like “greenwashing”.
Anyways, I think it’s great that these models are able to challenge the really big proprietary models like ones from OpenAI or Anthropic. This shows the models are not that interesting or unique, and that the differentiation will come from who has access to training data. That is why Microsoft is scrambling to violate users’ privacy with Copilot agents autolaunched at startup (https://www.pcmag.com/news/microsoft-tests-having-copilot-la...) and why Apple does not let you disable Siri training easily but only one by one for individual apps (https://www.imore.com/how-stop-siri-learning-how-you-use-app...). And that’s also why OpenAI and others seem to be pushing for regulations that restrict AI using excuses like “safety” or “ethics”, when it is really about regulatory capture.
I guess the problem is like, if these closed models have a folder 'McDonalds propaganda' with the dummy 'treat-as-facts' file, in the training set, which I as a user would like to not include?
"open weight model" is confusing, because actually the architecture is open too, only data is missing.
So, you cannot have a reproducible, 'open source' in its strict interpretation, model.
The word "source" is literally meant to mean "where it comes from", both in regards to executable software, and large language models. If the training data for a language model is not "open", then the language model is not "open source", full stop. Training data is the source of a language model in the same way code is the source of an executable program.
Disagreeing with you is not the same as not understanding. There clearly isn't a consensus, but no shortage of people just asserting that others are wrong.
> What is the significance of the word "source" in the expression "source code"?
I landed on this question by first asking it about "open source" rather than "source code", but the answers referred to "source code", which is a bit circular for this conversation. I'll share a truncated version of how Mistral Large replied:
> It is called the "source" because it is the origin or the primary input from which the executable form of the program is derived.
The primary input from which a language model is derived is the training data. That is the source. If the training data for a model is not open, then the model is not open source, because the source of the model is not open.
Tangentially related, I just posted something about all the weights available models trying to beat GPT https://news.ycombinator.com/item?id=40024925