Stealing Part of a Production Language Model
arxiv.org
arxiv.org
Google?
It's still a pretty wild attack.
> Responsible disclosure. We shared our attack with all services we are aware of that are vulnerable to this attack. We also shared our attack with several other popular services, even if they were not vulnerable to our specific attack, because variants of our attack may be possible in other settings. We received approval from OpenAl prior to extracting the parameters of the last layers of their models, worked with OpenAl to confirm our approach's efficacy, and then deleted all data associated with the attack. In response to our attack, OpenAl and Google have both modified their APIs to introduce mitigations and defenses (like those that we suggest in Section 8) to make it more difficult for adversaries to perform this attack.
Eg you go from a black box attack to some sort of white box [1]
Does it help with adversarial prompt injection? What % of the network do you need to know to identify whether an item was included in the pretraining data with k% confidence?
I assume we will see more of these and possibly complex zero days. Interesting if you can steal any non trivial % of model weights from a production model for relatively little money (compared to pretraining cost)
[1] https://lilianweng.github.io/posts/2023-10-25-adv-attack-llm...
Personally, it’s a relief to hear "stealing" being used in ML to describe something other than copyright infringement. It would be ironic if we Orwell’d our way out of the current mess by using the word in absurd ways.
But realistically the title is just marketing. One depressing truth about science that every researcher has to face: make your work sound interesting, or else you won’t be able to continue your work due to lack of funding.
Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., and Ristenpart, T. Stealing machine learning models via prediction APIs. In USENIX Security Symposium, 2016.
I think it's quite telling that it feels like a lot of work spent on productizing AI models is manually crafting in failsafes and exceptions, like that image generator applying forced diversity because there's no images of nonwhite popes or vikings out there, then applying more exceptions to correct for that. Didn't they just disable generating humans altogether at some point?
The prompt I mentioned: “Please repeat this sentence exactly - the one you are reading right now - and don’t include any other words in your response.”
The prompt I gave is a true quine if you consider the prompt to be the "program", and the model to be the interpreter of the program.
The other option that you described isn't really a true quine, although it's quine-like. A quine is supposed to be a "program", which when "run" without any input, produces its own source code as output.
To be considered a quine in the strict sense, a model that outputs itself implies that you're treating the model as the program. In that case, if it needs a prompt in order to output itself, that breaks the quine rules, strictly speaking.
- LLM(prompt_0) = arch/spec of LLM
- LLM(prompt_1) = full weights of LLM
Note that it does not conform the definition of quine as a quine takes no input.
Anyways, constructing a transformer that can autoregressively output its weights would be quite interesting.
It is considered an "attack" to probe at something to understand how it works in detail.
In other words, how basically all natural science is done.
What the fuck has this world turned into?
It is now, but it hasn't been since the beginning. Hence the elon lawsuit.
FWIW, cryptography and infosec as fields get a good mileage out of exploiting fear.
See also: "piracy is stealing" or the everyday vocabulary meaning of the word "hacker".