A computer program is a set of instructions that may be executed. Weights are values that may be loaded by a program, but are not a program in and of themselves.
A computer program is a set of instructions that may be executed. Weights are values that may be loaded by a program, but are not a program in and of themselves.
Would you consider a Python program to be data rather than program just because it is text input to the python interpreter instead of machine code for the CPU?
Weights are literally numbers computed as output. They are not instructions. The semantics of those numbers even when emplaced (loaded) in an artificial neural net is such that they do not execute. They are not instructions. LLM engines and diffusers perform searches where the weights are used to calculate additional output.
Is source code, like Python text, data? Yes. All code is data. But not all data are source code.
If I gave you a web request log, you would not assert it is a program. If I gave you a CSV file with time-series values from a sensor, you would not assert it is a program. If I hand you a database of contact information, you would not assert it is a program. Weight files are the equivalent of CSV files. They are are a dump of parameter values computed from training.
They are not a program.
The definition of computer program is well worn. So is the definition of source code, and the definition of parameters. Weights are parameters.
They are instructions if you consider the LLM system itself to be a kind of weird, indirect virtual machine. Each number can be mapped to a set of instructions that are executed. Even your CPU uses numbers (machine codes) to execute.
Join me in saying: ...code is data is code is data is code is data...
There is no real line between code and data. This is an observation that runs all the way from Turing Machines in computability theory to the Von Neumann architecture and homoiconicity in Lisp.
What we call 'data' is just code that needs a cleverer interpreter.
Code or data: Well... both.
If I told you the economy can be accurately modelled by
GDP(x) = Ax + B
But I don’t define A And B for you because it’s proprietary, you haven’t learned anything other than what you can glean from the structure of the model itself (it’s linear, there’s only a single input etc)
If most of these models are similarly structured, I’d say the weights are the program.
Parameters, or actual arguments, are values; data. Not instructions.
Valuable data is still data. It's significance doesn't magically turn it into source code.
At a minimum, it would be an active area of negotiation that the attorneys would take notice of. Source: have negotiated these agreements.
I imagine it is not settled law, but there's a clear argument to be made that regardless of the difficulty in curating the data set, it's still a data set.
Can it be licensed and sold. Yes, surely. Is it proper to pretend an open source license is sufficient protection, probably not.