Models are similar to code you might run on a compiled vm or native operating system. Llama.cpp is to a model as Python is to a python script. The license lays out the rights and responsibilities of the users of the software, or the model, in this case. The training data, process, pipeline to build the model in the first place is a distinct and separate thing from the models themselves. It'd be nice if those were open, too, but when dealing with just the model:
If it uses an OSI recognized open source license, it is an open source model. If it doesn't use an OSI recognized open source license, it's not.
Llama is not open source. It's corporate freeware.
It’s a bit like arguing that Linux is not open source because you don’t have every email Linus and the maintainers ever received. Or that you don’t know what lectures Linus attended or what books he’s read.
The weights “are the thing” in the same sense that the “code is the thing”. You can modify open code and recompile it. You can similarly modify weights with fine tuning or even architectural changes. You don’t need to go “back to the beginning” in the same sense that Linux would continue to be open source even without the Git history and the LKM mailing list.
Linux is open source, because you can actually compile it yourself! You don't need Linus's email for that (and if you needed some secret cryptographic key on Linus' laptop to decrypt and compile the kernel, then it wouldn't make sense to call it open-source either).
A language model isn't a piece of code, it's a huge binary blob that's being executed by a small piece of code that contains little of the added value, everything that matters is in the blob. Sharing only the compiled blob and the code to run makes it unsuitable for an “open source qualifier” (It's kind of the same thing as proprietary Java code: the VM is open-source but the bytecode you run on it isn't).
And yes, you can fine-tune and change things in the model weights themselves the same way you can edit the binary of a proprietary game to disable DRMs, that doesn't make it open-source either. Fine tuning doesn't give you the same level of control over the behavior of the model as the initial training does, like binary hacking doesn't give you the same control as having the source code to edit and rebuild.
In general, software rot is a huge issue, and many projects which may be of future archeological importance are increasingly non-reproducible as dependencies are often not vendored and checked into source, but instead downloaded at compile time from servers which lack strong guarantees about future availability.
Who were the countless unknown contemporaries of Giotto and Cimabue? Of Da Vinci and Michelangelo? Most of what we know about Renaissance art comes from 1 guy - Giorgio Vasari. We have more diverse information about ancient Egypt than the much more recent Italian Renaissance because of, essentially, better preservation techniques.
Compliance, interoperability, and publishing platforms for all this work (HuggingFace, Ollama, GitHub, HN) are our cathedrals and clay tablets. Who knows what works will fill the museums of tomorrow.
So, yeah, unless you own your own world-class datacenter, complete with the nuclear reactor necessary to power the training run, then training is not an option.
I think Zuck's discussion of energy being the limiting factor was one of the more interesting and surprising things to come out of the Dwarkesh interview. We're used to discussion of the $1B, $10B, $100B training runs becoming unsustainable, and chip shortages as an issue, but (to me at least!) it was interesting to see Zuck say that energy usage will be a disruptor before those do (partly because of lead times and regulations in expanding power supply, and bringing it in to new data centers). The sheer magnitude of projected power consumption needed is also interesting.
This is where I have to disagree. Continuing the training of an open model is the same process as the original training run. It's not a fundamentally different operation.
In practice it's not (because LoRA) but that doesn't matter: continuing the training is just a patch on top of the initial training, it doesn't matter if this patch is applied through gradient descent as well, you are completely dependent on how the previous training was done, and your ability to overwrite the model's behavior is limited.
For instance, Meta could backdoor the model with specially crafted group of rare tokens to which the model would respond a pre-determined response (say “This is Llama 3 from Meta” as some kind of watermark), and you'd have no way to figure out and get rid of it during fine-tuning. This kind of things does not happen when you have access to the sources.
That's one of many techniques, and is popular because it's cheap to implement. The training of a full model can be continued with full updates, the same as the original training run.
> completely dependent on how the previous training was done, and your ability to overwrite the model's behavior is limited.
Not necessarily. You can even alter the architecture! There have been many papers about various approaches such as extending token window sizes, or adding additional skip connections, quantization, sparsity, or whatever.
> specially crafted group of rare tokens
The analogy here is that some Linux kernel developer could have left a back door in the Linux kernel source. You're arguing that Linux would only be open source if you could personally go back to the time when it was an empty folder on Linus Torvald's computer and then reproduce every step it took to get to today's tarball of the source, including every Google search done, every book referenced, every email read, etc...
That's not what open source is. The code is open, not the process that it took to get there.
Linux development may have used information from copyrighted textbooks. The source code doesn't contain the text of those textbooks, and in some sense could not be "reproduced" without the copyrighted text.
Similarly, AIs are often trained on copyrighted textbooks but the end result is open source.
You can alter the architecture, but you're still playing with an opaque blob of binary *you don't know what it's made of*.
> The analogy here is that some Linux kernel developer could have left a back door in the Linux kernel source. You're arguing that Linux would only be open source if you could personally go back to the time when it was an empty folder on Linus Torvald's computer and then reproduce every step it took to get to today's tarball of the source, including every Google search done, every book referenced, every email read, etc...
No, it is just a bad analogy. To be sure that there's no backdoor in the Linux kernel, the code itself suffice. That doesn't mean there can be no backdoor since it's complex enough to hide things in it, but it's not the same thing as a backdoor hidden in a binary blob you cannot inspect even if you had a trillion dollar to spend on a million of developers.
> The code is open, not the process that it took to get there.
The code is by definition a part of a process that gets you a piece of software (which is the actually useful binary), and it's the part of the process that contains most of the value. Model weights are binary, and they are akin to the compiled binary of the software (training from data being a compute-intensive like compilation from source code, but orders of magnitude more intensive).
> Similarly, AIs are often trained on copyrighted textbooks but the end result is open source.
Court decisions are pending on the mere legality of such training, and it has nothing to do with being open-source, what's at stake is whether or not these models can be open-weight or if it is copyright infringement to publish the models.
There are places (e.g. in the Linux kernel? AMD drivers?) where lots of generated code is pushed and (apart from the rants of huge unwieldy commits and complaints that it would be better engineering-wise to get their hands on the code generator, it seems no one is saying the AMD drivers aren't GPL compliant or OSI-compliant?
There are probably lots of OSS that is filled with constants and code they probably couldn't rederive easily, and we still call them OSS?
The starting point is the ability to run the LLM as you wish, for any purpose - so if a license prohibits some uses and you have to start any usage with thinking whether it's permitted or not, that's a fail.
Then the freedom where "source" matters is the practical freedom to change the behavior so it does your computing as you wish. And that's a bit tricky - since one interpretation would require having the training data, training code and parameters; but for current LLMs the training hardware and cost of running it is a major practical limitation, so much that one could argue that the ability to change the behavior (which is the core freedom that we'd like) is separate from the ability to recreate the model, and would be more relevant in the context of the "instruction training" which happens after the main training, is the main determiner of behavior (as opposed to capability), and so the main "source would be the data for that (instruct training data, and the model weights before that finetuning) so that you can fine-tune the model on different instructions, which requires much less resources than training it from scratch, and don't have to start with the instructions and values imposed on the LLM by someone else.
Most of these other models, like Llama, are open weight not open source - and open weight is just openwashing, since you’re just getting the final output like a compiled executable. But even with OLMo (and others like Databrick’s DBRX) there are issues with proprietary licenses being used for some things, which prevent truly free use. For some reason in the AI world there is heavy resistance to using OSI-approved licenses like Apache or MIT.
Finally, there is still a lack of openness and transparency on the training data sets even with models that release those data sets. This is because they do a lot of filtering to produce those data sets that happen without any transparency. For example AI2’s OLMo uses a dataset that has been filtered to remove “toxic” content or “hateful” content, with input from “ethics experts” - and this is of course a key input into the overall model that can heavily bias its performance, accuracy, and neutrality.
Unfortunately, there is a lot missing from the current AI landscape as far as openness.
seems like they make everything available.