You can also edit binaries by hand.
Especially with dynamically linked binaries like many games.
Seriously, a set of weights that already works really well is basically the ideal basis for a _lot_ of ML tasks.
That's fundamentally the same thing though - you run an optimization algorithm on a binary blob. I don't see why this couldn't work. Sure, a neural net is designed to be differentiable, while ELF and PE executables aren't, but then backprop isn't the be-all, end-all of optimization algorithms.
Off the top of my head, you could reframe the task as a special kind of genetic programming problem, one that starts with a large program instead of starting from scratch, and that works on an assembly instead of an abstract syntax tree. Hell, you could first decompile the executable and then have the genetic programming solver run on decompiled code.
I'd be really surprised if no one tried that before. Or, if such functionality isn't already available in some RE tools (or as a plugin for one). My own hands-on experience with reverse engineering is limited to a few attempts at adding extra UI and functionality to StarCraft by writing some assembly, turning it into object code, and injecting it straight into the running game process[0] - but that was me doing exactly what you described, just by hand. I imagine doing such things is common practice in RE that someone already automated finding the specific parts of the binary that produce the outputs you want to modify.
--
[0] - I sometimes miss the times before Data Execution Prevention became a thing.
"Given specific inputs X and outputs Y, have a computer automatically find modifications to F so that F(X) gives Y" is a problem that's been studied for nearly a century now (longer, if relax the meaning of "computer"), with plenty of well-known solutions, most of which don't require F to be differentiable.
Isn't "operational research" a standard part of undergrad CS curriculum? It was at my alma mater.
This is like saying, hey, a regular binary executable is fine because I can edit it with hexl-mode.
Or should I say, it held water until few days ago.
Personally though, I never bought it. Saying that weights are the "preferred form of the work for making modifications to it" because a) approximately no one can afford to start with the training data, and b) fine-tuning and training LoRAs are cheap enough, is basically like saying binary blobs are "open source" as long as they provide an API (or ABI) for other programs to use. By this line of reasoning, NVIDIA GPU stack and Broadcom chipset firmware would qualify as open source, too.
I don't build my browser, it's too expensive, but the cost of building has nothing to say with how open the access to things. It'd be cool if the community could fork the project, propose changes and maybe crowdfund a training/build run to experiment.
But probably it is impossible for them to release the training data, as they have probably not made it all reproducible, but live ingested the data, and the data has since then chanced in many places. So the code to live ingest the data becomes the actual source, I guess.
I think a closer comparison would be Android and GApps, where if you remove the latter, most would deem the phone unusable.
if open the source then it is open source
if you write a book/blog about how you came up with the ideas but didn't publish the source it's not open source, even if you publish the blog+binaries
If the magic values are some kind of microcode or firmware, or something else that is executed in some way, then no, it is not really open source.
Even algorithms can be open source in spirit but closed source in practice. See ECDSA. The NSA has never revealed in any verifiable way how they came up with the specific curves used in the algorithm, so there is room for doubt that they weren't specifically chosen due to some inherent (but hard to find) weakness.
I don't know a ton about AI, but I gather there are lots of areas in the process of producing a model where they can claim everything is "open source" as a marketing gimmick but in reality, there is no explanation for how certain results were achieved. (Trade secrets, in other words.)
To my understanding, the contents of a .safetensors file is purely numerical weights - used by the model defined in MIT-licensed code[0] and described in a technical report[1]. The weights are arguably only really "executed" to the same extent kernel weights of a gaussian blur filter would be, though there is a large difference in scale and effect.
[0]: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/inferen...
- Windows MetaFiles (WMF, EMF, EMF+), still in use (mostly inside MS Office suite) - you'd think they're just another vector image format, i.e. clearly "data", but this one is basically a list of function calls to Windows GDI APIs, i.e. interpreted code.
- Any sufficiently complex XML or JSON config file ends up turning into an ad-hoc Lisp language, with ugly syntax and a parser that's a bug-ridden, slow implementation of a Lisp runtime. People don't realize that the moment they add conditionals and ability to include or refer back to other parts of config, they're more than halfway to a Turing-complete language.
- From the POV of hardware, all native code is executed "to the same extent kernel weighs of a gaussian blur filter" are. In general, all code is just data for the runtime that executes it.
And so on.
Point being, what is code and what is data depends on practical reasons you have to make this distinction in the first place. IMHO, for OSS licensing, when considering the reasons those licenses exist, LLM weights are code.
Another note is that this may be the more ethical option. I’m sure the training data contained lots of copyrighted content, and if my content was in there I would prefer that it was released as opaque weights rather than published in a zip file for anyone to read for free.
IMO it should be considered freeware, and only partially open. It's like releasing an open source program with a part of it delivered as a binary.
- open-source inference code
- open weights (for inference and fine-tuning)
- open pretraining recipe (code + data)
- open fine-tuning recipe (code + data)
Very few entities publish the later two items (https://huggingface.co/blog/smollm and https://allenai.org/olmo come to mind). Arguably, publishing curated large scale pretraining data is very costly but publishing code to automatically curate pretraining data from uncurated sources is already very valuable.
It is a bit more permissive than Llama's it seems (no MAU threshold it seems).
But the recent DeepSeek-R1-Zero and DeepSeek-R1 have MIT licensed weights.
https://huggingface.co/deepseek-ai/DeepSeek-V3/tree/main https://huggingface.co/deepseek-ai/DeepSeek-R1/tree/main
So when we apply the same principles to another category, such as weights, we should not call things “open” that don’t grant those same freedoms. In the case of this research license, Freedom 0 at least is not maintained. Therefore, the weights aren’t open, and to call them “open” would be to indeed dilute the meaning of open qua open source.
So it seems to me that it's at least dubious if those restricted licences can be enforced (that said you likely need deep pockets to defend yourself from a lawsuit)
Not necessarily[0], it's a WIP, but: https://github.com/huggingface/open-r1
[0] Surely they won't end up with the exact same weights, but it should be possible to verify something about the model and approach
Yes training is left as an exercise to the user, but it's outlined in the paper, and a good ML engineer should be able to get started with it, cluster of GPUs not included
"3.2.2. Efficient Implementation of Cross-Node All-to-All Communication
In order to ensure sufficient computational performance for DualPipe, we customize efficient cross-node all-to-all communication kernels (including dispatching and combining) to conserve the number of SMs dedicated to communication. The implementation of the kernels is codesigned with the MoE gating algorithm and the network topology of our cluster. To be specific, in our cluster, cross-node GPUs are fully interconnected with IB, and intra-node communications are handled via NVLink. NVLink offers a bandwidth of 160 GB/s, roughly 3.2 times that of IB (50 GB/s). To effectively leverage the different bandwidths of IB and NVLink, we limit each token to be dispatched to at most 4 nodes, thereby reducing IB traffic. For each token, when its routing decision is made, it will first be transmitted via IB to the GPUs with the same in-node index on its target nodes. Once it reaches the target nodes, we will endeavor to ensure that it is instantaneously forwarded via NVLink to specific GPUs that host their target experts, without being blocked by subsequently arriving tokens. In this way, communications via IB and NVLink are fully overlapped, and each token can efficiently select an average of 3.2 experts per node without incurring additional overhead from NVLink. This implies that, although DeepSeek-V3 13 selects only 8 routed experts in practice, it can scale up this number to a maximum of 13 experts (4 nodes × 3.2 experts/node) while preserving the same communication cost. Overall, under such a communication strategy, only 20 SMs are sufficient to fully utilize the bandwidths of IB and NVLink.
In detail, we employ the warp specialization technique (Bauer et al., 2014) and partition 20 SMs into 10 communication channels. During the dispatching process, (1) IB sending, (2) IB-to-NVLink forwarding, and (3) NVLink receiving are handled by respective warps. The number of warps allocated to each communication task is dynamically adjusted according to the actual workload across all SMs. Similarly, during the combining process, (1) NVLink sending, (2) NVLink-to-IB forwarding and accumulation, and (3) IB receiving and accumulation are also handled by dynamically adjusted warps. In addition, both dispatching and combining kernels overlap with the computation stream, so we also consider their impact on other SM computation kernels. Specifically, we employ customized PTX (Parallel Thread Execution) instructions and auto-tune the communication chunk size, which significantly reduces the use of the L2 cache and the interference to other SMs."
It's definitely not the full model written in PTX or anything, but still some significant engineering effort to replicate, from people commanding 7-figure salaries in this wave, since the training code isn't open.
I wish Intel didn't kill the maxas effort back then (buying the team out and ... what ?) as they were going even lower down the stack.
It’s not and I called it [1].
We had three options: (A) Open weights (favoured by Altman et al); (B) Open training data (favoured by some FOSS advocates); and (C) Open weights and model, which doesn’t provide the training data, but would let you derive the weights if you had it.
OSI settled on (C) [2], but it did so late. FOSS argued for (B), but it’s impractical. So the world, for a while, had a choice between impractical (B) and the useful-if-flawed (A). The public, predictably, went with the pragmatic.
This was Betamax vs VHS, except in natural linguistics. There is still hope for (C). But it relies on (A) being rendered impractical. Unfortunately, the path to that flows through institutionalising OpenAI et al’s TOS-based fair use paradigm. Which means while we may get a definition (not exactly (B), but (A) absent use restrictions) we’ll also get restrictions on even using Chinese AI.
If you want to prioritize pragmatism, that every discussion of this includes a lengthy "so what open source do you mean, exactly?" subthread proves this was a poor choice. It causes uncertainly that also makes it harder for the folks releasing these models to make their case and be taken seriously for their approach.
We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture" to appreciate the Python file that utilizes the weights more.
Technically yes, practically no.
You’re describing a prisoner’s dilemma. The term was available, there was (and remains) genuine ambiguity over what it meant in this context, and there are first-mover advantages in branding. (Exhibit A: how we label charges).
> causing collateral damage outside the AI bubble, and is nothing like Betamax vs. VHS
Standards wars have collateral damage.
> We should probably call them "free to run", if the "it's cheap" connotation of "freeware" needs to be avoided. Or maybe "open architecture"
Language is parsimonious. A neologism will never win when a semantic shift will do.
Agreed, but I think it's worth lamenting the danger in that. History is certainly full of transitory calamity and harm when semantic shifts detach labels from reality.
I guess we're in any case in "damage is done" territory. The question is more about where to go next. It does appear that the term "open source" isn't working for what these folks are doing (you could even argue whether the "available" term they chose was a strong one to lean on in the first place), so we'll see what direction the next shift takes.
Sort of. We can learn from the example. Perfect is the enemy of the good.
It’s ambiguously open.
Then it’s never distributable and any definition of open source requiring it to be is DOA. It’s interesting, as an argument against copyright. But that academic.
Also I would say arguing that the model weights are just the "binary" is disingenuous, because nobody wants releases that only contain the training data and scripts to train and not the model weights (which would be perfectly fine for open source software if we argue that the weights are just the binaries), because they would be useless to almost everyone, because they don't have the resources to train the model.
I agree it would nice to know the details of their training, but, simply calling this drop an "opaque binary" is seriously underselling it no?
https://github.com/deepseek-ai
More importantly, they spelled out their methodology in depth in a paper (the code/implementation is trivial in comparison to the methodology)
I'm very appreciative of what they've done, but it's open weights and methodology, not open source.
Paxos isn’t open source just because you can read the paxos paper.
But I was merely adding the missing context using the (sorry) lingua Franca of AI.
Things like tweaking all the hyperparameters to make the training process actually work may be more tricky though.
It's more akin to scientific research where everyone is using the same molecules, but depending on the process you put the molecules through, you get a different outcome.
How is that different than the 0s and 1s of a program?
Assembly instructions are literally standard. What’s more, if said program uses something like Java, the byte code is even _more_ understandable. So much so that there is an ecosystem of Java decompilers.
Binary files are not the “source” in question when talking about “open source”
Google Spanner has a nice white paper but you wouldn't consider it open source, for example.
tldr: we're already on to the next model, don't expect anything else to get open sourced.
> I was just told that the amount of people there are too limited, and open-sourcing needs another layer of hard work beyond making the training framework brrr on their own infra. So their priority has been to open-source everything that is MINIMUM + NECESSARY to the community while pushing most efforts on iterating to the next generation of models I think. They have been write everything clearly in technical reports and encourage the community to engage in reproduction , which is the unique insight of the team as well I think.
But China did distribute them with sharing-friendly terms, what is completely different from others, like Meta, and makes the name way less misleading this time.