While the training code and data are the true source. Since if you want to robustly modify the LLM that's actually what you need.
But since "compilation" (training) is extremely compute intensive this isn't something accessible to anyone without an entire datacenter.
Anyway semantics aside having the binary is still infinitely better than dealing with an api as far as privacy and control go.
I don't know LLM theory well enough to say if there's some secret sauce they can hold back that makes training ineffective. Less effective I'm sure, we don't have access to their smart training schemes, but post-training should always be possible IIUC.
post-training is like writing a wrapper around the binary. It is closer to building on top of than truly modifying, in that you can tailor things to your needs slightly but cannot make fundamental changes to the underlying thing.
For a stretched analogy, I think it is more like LEGO sets. Someone hands you a 10,000 piece masterpiece, and a box of unused LEGO parts. Hackers on HN object that the LEGO part manufacturing process is not included, you can't make your own parts, etc. But it's LEGO. You can pull apart the model, see how it is constructed, add your own refinements and features, or even redo it from the ground up. In a practical sense having knowledge about the factory making the parts doesn't really matter here.
https://en.wikipedia.org/wiki/Ablation_(artificial_intellige...
Put in the work. This is akin to asking how to remove Rust from a Rust project; just because something is legally available to you doesn't mean you wont need to apply dome elbow grease, depending how deep the changes you want are, ablation, fine-tuning, or distillation are tools you can use to remove "censorship"
Open source means you reveal how you created this binary.
you need to "literally" go read the definition of open source software or even ask an LLM to define it for you. Weights + inference code are not the source code they're more like the compiled binary. Making modifications to the behavior of a model with additional training is like writing a mod for minecraft. Sure, you can change things but it doesn't make it open source.
Calling these models "open source" is an old trap that software companies use to use. Free to download but then, once you're fully comitted, the trap snaps shut and you must pay up to continue.
The actual problem is that we know nothing about the training set of any open-weights model. They could be intentionally biased to influence users, from political censorship to brand advertising, or general shaping of cultural norms. You run the model on own hardware not knowing if it is designed to act against you. Having whole chain open source would allow audit and reproducing the results.
Training data and code.