I remember the NVIDIA Linux kernel binary blob driver discussions from the early-mid 2000s. Who knew we had an open source driver all along...
Open-source as a concept doesn’t really correspond well with LLMs but to the extent that it does, access to the training data is not required because that training data is not the preferred form for making modifications.
I definitely disagree with this.
Yes, you can do some SFT fine tuning on an existing model, but if you want to make specific, substantial, targeted changes (less safety? better performance on math and code at the expense of general knowledge?), your best bet is to change the training mixture, and for that you need the original datasets.
linus turns to the camera, giving a thumbs up