If the FOSS community sets this as the benchmark for open source in respect of AI, they're going to lose control of the term. In most jurisdictions it would be illegal for the likes of Meta to release training data.
If the FOSS community sets this as the benchmark for open source in respect of AI, they're going to lose control of the term. In most jurisdictions it would be illegal for the likes of Meta to release training data.
Please read through their "acceptable use" policy before you decide whether this is really in line with open source.
I'm not taking a specific posiion on this license. I haven't read it closely. My broad point is simply that open source AI, as a term, cannot practically require the training data be made available.
How come releasing an LLM trained on that data is not illegal then? I think it should be.
Sure. But that's not going to be released. The term open source AI cannot be expected to cover it because it's not practical.
Open training data is hard to the point of impracticality. It requires excluding private and proprietary data.
Meanwhile, the term "open source" is massively popular. So it will get used. The question is how.
Meta et al would love for the choice to be between, on one hand, open weights only, and, on the other hand, open training data, because the latter is impractical. That dichotomy guarantees that when someone says open source AI they'll mean open weights. (The way open source software, today, generally means source available, not FOSS.)
you are playing very loosely with terms that have specific, widely accepted definitions (e.g. https://opensource.org/osd )
I don't get why you think it would be useful to call LLMs with published weights "open source"
OSF's definition is far from the only one [1]. Switzerland is currently implementing CH Open's definition, the EU another one, et cetera.
> I don't get why you think it would be useful to call LLMs with published weights "open source"
I don't. I'm saying that if the choice is between open weights or open weights + open training data, open weights will win because the useful definition will outcompete the pristine one in a public context.
[1] https://en.wikipedia.org/wiki/Open-source_software#Definitio...
For the CH Open, I'm not finding anything specific, even from Swiss websites, could you help me understand what you're referring to here?
I'm guessing that all these definitions have at least some points in common, which involves (another guess) at least being able to produce the output artifacts/binaries by yourself, something that you cannot do with Llama, just as an example.
Was on the HN front page earlier [1][2]. The definition comes strikingly close to source on request with no use restrictions.
> all these definitions have at least some points in common
Agreed. But they're all different. There isn't an accepted defintiion of open source even when it comes to software; there is an accepted set of broad principles.
[1] https://news.ycombinator.com/item?id=41047172
[2] https://joinup.ec.europa.eu/collection/open-source-observato...
Agreed, but are we splitting hairs here and is it relevant to the claim made earlier?
> (The way open source software, today, generally means source available, not FOSS.)
Do any of these principles or definitions from these orgs agree/disagree with that?
My hypothesis is that they generally would go against that belief and instead argue that open source is different from source available. But I haven't looked specifically to confirm if that's true or not, just a guess.
I don't think so. Take the Swiss definition. Source on request, not even available. Yet being branded and accepted as open source.
(To be clear, the Swiss example favours FOSS. But it also permits source on request and bundles them together under the same label.)
Realistically, nobody outside of Hacker News commenters have ever cared about the OSD. It's just not how the term is used colloquially.
and (strong personal opinion) any software developer should have a firm grip on the terminology and details for legal reasons
There is a large span of people between gray beard programmer and lay person, and many in that span have some concept of open-source. It's often used synonymously with visible source, free software, or in this case, open weights.
It seems unfortunate - though expected - that over half of the comments in this thread are debating the OSD for the umpeenth time instead of discussing the actual model release or accompanying news posts. Meanwhile communities like /r/LocalLlama are going hog wild with this release and already seeing what it can do.
> any software developer should have a firm grip on the terminology and details for legal reasons
They'd simply need to review the terms of the license to see if it fits their usage. It doesn't really matter if the license satisfies the OSD or not.
Right, so the onus is on Facebook/Meta to get that right, then they could call something Open Source, until then, find another name that already doesn't have a specific meaning.
> (The way open source software, today, generally means source available, not FOSS.)
No, but it's going in that way. Open Source, today, still means that the things you need to build a project, is publicly available for you to download and run on your own machine, granted you have the means to do so. What you're thinking of is literally called "Source Available" which is very different from "Open Source".
The intent of Open Source is for people to be able to reproduce the work themselves, with modifications if they want to. Is that something you can do today with the various Llama models? No, because one core part of the projects "source code" (what you need to reproduce it from scratch), the training data, is being held back and kept private.
Here's the source of the disagreement. You're justifying the use of the term "open source" by saying it's logical for Meta to want to use it for its popularity and layman (incorrect) understanding.
Other person is saying it doesn't matter how convenient it is or how much Meta wants to use it, that the term "open source" is misleading for a product where the "source" is the training data, and the final product has onerous restrictions on use.
This would be like Adobe giving Photoshop away for free, but for personal use only and not for making ads for Adobe's competitors. Sure, Adobe likes it and most users may be fine with it, but it isn't open source.
>The way open source software, today, generally means source available, not FOSS.
I don't agree with that. When a company says "open source" but it's not free, the tech community is quick to call it "source available" or "open core".
I'm actually not a fan of Meta's definition. I'm arguing specifically against an unrealistic definition, because for practical purposes that cedes the term to Meta.
> the term "open source" is misleading for a product where the "source" is the training data, and the final product has onerous restrictions on use
Agree. I think the focus should be on the use restrictions.
> When a company says "open source" but it's not free, the tech community is quick to call it "source available" or "open core"
This isn't consistently applied. It's why we have the free vs open vs FOSS fracture.
Who? It's not their data.
Synthetic part of the training data could be released.
For an LLM, that’s not the training data. That’s the model itself. You don’t make changes to an LLM by going back to the training data and making changes to it, then re-running the training. You update the model itself with more training data.
You can’t even use the training code and original training data to reproduce the existing model. A lot of it is non-deterministic, so you’ll get different results each time anyway.
Another complication is that the object code for normal software is a clear derivative work of the source code. It’s a direct translation from one form to another. This isn’t the case with LLMs and their training data. The models learn from it, but they aren’t simply an alternative form of it. I don’t think you can describe an LLM as a derivative work of its training data. It learns from it, it isn’t a copy of it. This is mostly the reason why distributing training data is infeasible – the model’s creator may not have the license to do so.
Would it be extremely useful to have the original training data? Definitely. Is distributing it the same as distributing source code for normal software? I don’t think so.
I think new terminology is needed for open AI models. We can’t simply re-use what works for human-editable code because it’s a fundamentally different type of thing with different technical and legal constraints.
I disagree with the purists - if you can legally change the source or weights - even without having access to the data used by the upstream authors - it's open enough for me. YMMV.