(As an aside, this is why the "open weights are not open source" thing is a complete misunderstanding. The weights themselves along with the documentation give you enough to fine tune the LLM. You can't rebuild it from scratch, but you can't do this even with the data anyway (because of randomness!))
> You can't rebuild it from scratch
There's is extremely clear and misunderstanding-free.
Open weights is not open source.
If you want to train it from scratch you need data, yes. But presumably if you are doing that there is a reason you want to do it.
You lose nothing without access to the original data - you can do every single modification without it.
That is unlike open source where you (mostly) need to source code to modify it beyond what the original designed originally thought.
Though in practice you can get all the benefits of both determinism and (that kind of) randomisation by using a PRNG and saving the seed you are using.
It's an open question roughly on par with P vs NP whether true randomisation is ever necessary, or whether PRNGs are enough. So far we haven't found any problem or algorithm where true RNG is necessary and good PRNG ain't enough.
I don't see the connection? Most local builds are done with debugging on and optimisation turned off anyway.
And what you deliver to your customers is usually something you produce on your CI/CD server, not what's on any developer's machine.
(And if you want reproducible builds https://en.wikipedia.org/wiki/Reproducible_builds you can't optimise for a specific wall clock time.)
> Also, most (optimizing) compilers ”optimize” the code for a fixed amount of time, leading to better optimized binaries on faster computers.
That’s why you should give your developers computers that have slow clocks!
They could give you the random seeds? (Assuming you carefully train in such a way to remove other sources of randomness, like concurrent execution.)
Though in principle saving random seeds is a lot less hassle than keeping entire checkpoints around: your random seeds would fit on a floppy disk or even a tweet. The checkpoint is basically as big as the model.
> It's entirely reproducible from the available documentation
You have a very interesting understanding of "reproducibility", I'll give you that :)
But even with that, there are plenty of technical details (especially in regards to the training process) missing from the tech report that leads to these weights not being reproducible in any sense of that word.
He was talking about two different things, hence the parentheses. The architecture is reproducible, not the model weights.
But what's the point of even saying that? Of course it is, otherwise how is it supposed to run in the runtimes? You cannot release model weights that others can run, without also releasing the model architecture, it's in the code at the very least...
The K3 arch can be implemented from the spec.
A PDF reader can not be implemented from the spec.
Have you ever tried reproducing even a small neural network exactly if you train on GPUs on more than one machine? I have and it is pretty close to impossible, and I'd argue actually impossible at scale.
Reproducing the exact training run, however, is basically impossible without the original dataset and training pipeline (here meaning all of the code + infra involved in actually executing the pre and post training loops). Also, it would be exorbitantly expensive to do if you weren't also a lab trying to train a similar model.
But you can still scale the architecture down and experiment as a solo researcher using the published research. There are probably some open source implementations already on GitHub for any given big open model release.
I share this expectation. But this is an interesting empiric question that deserves study; even if just to confirm what 'everyone knows'.
Heck even ffmpeg introduce randomness when stitching together downloaded chunks from youtube. By design.
Transformers are very "mendable" in that you can permute the architecture in crazy or random ways, and still basically always end up with get a coherent LLM. The difference comes down to training efficiency, inference efficiency, and usually minor differences in performance.
Hyperparams and stuff, I mean it's standard to do a sweep anyway.