linus turns to the camera, giving a thumbs up
Open-source as a concept doesn’t really correspond well with LLMs but to the extent that it does, access to the training data is not required because that training data is not the preferred form for making modifications.
I definitely disagree with this.
Yes, you can do some SFT fine tuning on an existing model, but if you want to make specific, substantial, targeted changes (less safety? better performance on math and code at the expense of general knowledge?), your best bet is to change the training mixture, and for that you need the original datasets.
But I agree, it's a real shame.
Literally how language has always worked and evolved, though.
Oof, I know there's a bunch of linguists and grammarians who are going to mock you for that bracket.
It is but it was the "correct" part attached to prescriptivism they'd be mocking because that is not how linguists and grammarians work (they are descriptivists and fond of making fun of prescriptivists.)
But this thread is about misuse of the term as applied to the weights package. Those of us who know what open source means should not continue to dilute the term by calling these LLMs by that term.
But we don't actually know all that much about how language really works, for all the resources we spend on linguistics - as the old IBM joke about AI goes, "quality of the product increases every time we fire a linguist" (which is to say, we consistently get better results by throwing "every written word known to man" at a blank model than we do by trying to construct things from our understanding).
All that said, just because we're taking a different, and quite possibly slower / less compute-efficient route, doesn't mean that we can't get to AGI in this way.
No, we can’t few shot it and we don't get there faster (but we develop a lot of other capabilities on the way.) We train on a lot more data; the human brain, unlike an LLM, is training on all that data in processes for ”inference”, and it receives sensory data estimated on the order of a billion bits per second, which means by the time we start using language we’ve trained on a lot of data (the 15 trillion tokens from a ~17 bit token vocabulary that Llama3 is something like the size of a few days of human sense data.) Humans just are trained on and process vastly richer multimodal data instead of text streams.
Yeah, humans don't acquire language separately from other experience.
> Most of the data that you reference is visual input and other body sensations that aren't directly related to that.
Visual input and other body sensations are not unrelated to language acquisition.
> OTOH humans don't take all that much text to learn to read and write.
That generally occurs well after they have acquired both language and recognizing and using symbolic visual communication, and they usually have considerable other input in learning how to read and write besides text they are presented with (e.g., someone else reading words out loud to them.)
It's just like even for a true open source software you still need to bring your own hardware to run it on.
AI2 has a model called OLMo that is actually open source. They share the training data, training source code, and many other things:
https://allenai.org/blog/olmo2
They also released an app recently, to do local inference on your phone with a small truly open source model:
It's not like they understand what the weights mean either and if they released the code and dataset used to create it, you probably couldn't recreate it, owning the fact that you don't own tens of thousands of GPUs.
If a software's source is released without all the documentation, commit history, bug tracker data etc., it's still considered open source, yet you couldn't recreate it without that information.
That’s like a chef giving you chicken instead of beef and calling it vegetarian.
A truly open model has open code that gathers pre-training data, open pre-training data, open RLHF data, open RLAIF data generated from its open constitution and so on.
The binary blob is the last thing I'd want - as a heavy user of LLMs I'm actually more interested in the detail of what all training data is in full, than I am the binary blob.
I see both sides here, but I don't think it's a hill worth dying on. The 'open source' part in this case is just not currently easily modifyable. That may not always be the case.
I think the two plausible answers are:
1. The person prompting (for example telling chatgpt 'please produce a fizzbuzz program') owns the copyright. The creativity lies in the prompt, and the chatgpt transformation is not transformative or meaningful.
2. The output of ChatGPT is derivative of the training data, and so the copyright is owned by all of the copyright holders of the input training data, i.e. everyone, and it's a glowing radioactive bomb of code in terms of copyright that cannot be used or licensed meaningfully in open source terms.
There are existing things like 1, where for example if someone takes a picture, and then uses photoshop to edit it, possibly with the "AI erase" tool thingy, they still own the photo's copyright. Photoshop transformed their prompt (a photo), but adobe doesn't get any copyright, nor do any of the test files adobe used to create their AI tool.
I don't think AI is like that, but it hasn't gone to court as far as I know, so no one really knows.
What do you think an open source matrix should look like?
I'm not even necessarily advocating that these things should be released, but the term "open source" has a pretty well-understood meaning that is being equivocated here.
Its about reproducibility and modifiability. Compiled executables (and their licences) lack that. The same as these downloadable blobs.
You can absolutely have open source machine code.
The issue is and always has been that you need to have access to the same level of abstraction as the people writing the source code. The GPL specifically bans transpilers as a way to get around this.
In ML there is _no_ level of abstraction other than the raw weights. Everything else is support machinery no different to an compiler, and os, or a physical computer to run the code on.
Linux isn't closed source because they don't ship a C compiler with their code. Why should llama models be any different?