Planting Undetectable Backdoors in Machine Learning Models
arxiv.org
arxiv.org
I can imagine it would be very very difficult to reverse engineer from the model that this training is there, and also very difficult to detect with testing. How would you know to test this particular case? The same could be done for many other models.
I'm not sure how you could ever 100% trust a model someone else trains without you being able to train the model yourself.
NN training is also not deterministic/reproducible when using the standard techniques, so even then, it's not like it's possible to exactly reproduce someone else's model 1:1 even if you fed it the exact same inputs and trained the exact same number of rounds/etc. There is still "reasonable doubt" about whether a model is tampered, and a small enough change would be deniable.
(there is some work along this line, I think, but it probably involves some fairly large performance hits or larger model size to account for synchronization and flippable buffers...)
It's generally considered that those variations should not impact model accuracy (other mentioned concerns like randomization for initialization, dropout or sample selection do affect accuracy so there are tools to ensure that they are reproducible from the random seeds), and we care a lot about training performance unless model accuracy is impacted, so there's not much engineering attention paid to ensuring that the model weights would be exactly identical and verifiable, most users would not accept a performance hit for that.
I've done lots of ensembling work where we train multiple copies of the model, and generally we would start with different seed each time. If we start with the same seed but don't force the training to be deterministic, the results are typically different on each training run, though I have not actually explored if they are "less different" than if you start with different random seeds for initializing everything. There is that loss landscape paper that looks at how the weights vary for different kinds of perturbations, it would be interesting to try the same thing with gpu thread noise as the only source of randomness and see what happens
Last few days I've been noticing alot of ego filled arguing, or maybe I've been spending too much time on hn.
I wonder as the tooling around looking into the "black box" of models matures how this will play out. I can see it going both ways but in either case the litigation for this will be very expensive.
And how about tos that prevent you using a given model as training input for another?
Are sure you can even trust the models you train yourself? It's possible that a model you trained is defective in a way you don't realize, e.g. will not recognize speeding cars with a sticker that has a picture of a bicycle[1]. It's likely that someone will discover that vulnerability before the publisher does, given the current state of ML model observability; ML model exploits are going to be wild, and inexplicable.
1. or the letters "GTEHQ"
The hardest part of this problem is the difficulty of auditing. If you use someones open source code you at least have to potential of reading the code to look for something... with a large model that is difficult on a different scale!
It's unlikely your test data would contain such a picture. Somebody else can notice this loophole and abuse it.
Applying this to a practical ML model is of course left as an exercise for the reader. While the research certainly proves that it's fundamentally possible (and mathematically trivial) to do such a thing, I feel that the structures of ML models are relatively transparent in most practical applications, making it comparatively easy to detect "parallel" verification networks thus constructed. The dataflow graph will be pretty revealing, but the victim would have to actually inspect it in the first place.
Of course, that in turn makes it a game of obfuscation - can you inconspicuously hide the signature check and final muxing step among the main network? I have no doubts that you can find a way if you're so determined.
But I think the most salient points of the paper are that 1. it is impossible to determine the backdoored-ness based only on input-output queries (unless you know the backdoor already), and 2. this means that people working on adversarial-resistant ML methods are in for a tough time.
There's more to be found in the paper, this is just my short summary after reading the most interesting bits.
When I first encountered the notion of adversarial examples, I thought it was a niche concern. As this paper outlined, however, the growth of "machine-learning-as-a-service" companies (Amazon, Open AI, Microsoft, etc.) has actually rendered this a legitimate concern. From my skimming, I wanted to highlight their interesting point that "gradient-based post-processing may be limited" in mitigating a compromised model. These points really bring these concerns from an academic to business realm.
Lastly, I'm delighted that they acknowledge their influences from the cryptographic community with respect to rigorously quantifying notions of "hardness" and "indistinguishable." Of note, they seem to base their undetectable backdoors on the assumption that the shortest vector problem is not in BQP. As I recently learned looking at the NIST post-quantum debacle, this has been a point of great contention.
I've in all likelihood mischaracterized the paper, but I look forward to reading it!
With respect to the shortest vector problem (SVP) being a point of contention among NIST PQC participants, two of the round 3 finalists are based on lattice cryptography, with NTRU directly relying on the hardness of SVP. The two concerns are:
1. The risks of lattice-based cryptography are poorly understood [1], [2]
2. Research progress into attacks on lattice-based cryptography have been fruitful during the NIST PQC process [1], [3].
From what I've gathered as a layperson, much of these concerns have been voiced by Daniel J. Bernstein. Bernstein contributed to the NTRU Prime software [4], which was used in OpenSSH 9 (I'll circle back to this point). As a consequence of these two concerns, the main argument seems to be that NIST should at least provide warnings [6] on the risks of lattice cryptography, particularly with regard to the use of cyclotomics by one of the finalists [5].
A common thread amongst these criticisms seems to be a distrust of NIST guidelines (a point that is also echoed by this ML backdoor paper). This has evidently stirred some bad blood between NIST workers and Bernstein [7], [8]. I'm sure to there's more to the story (especially since Bernstein's NTRU prime was a NIST PQC candidate), but I suppose NIST isn't free from passive-aggressiveness?
Within the context of this bad-blood, it's amusing that OpenSSH 9 uses Bernstein's NTRU Prime (doesn't use cyclotomics iirc), as opposed to one of NIST PQC's finalists.
(DISCLAIMER: I'm a layperson, and I encourage people to read the sources themselves to make an informed opinion. People are welcome to correct. )
[1] - See the link to the "Risks of lattice KEMs" PDF at the top: https://ntruprime.cr.yp.to/warnings.html
[2] - https://groups.google.com/a/list.nist.gov/g/pqc-forum/c/Fm4c...
[3] - https://groups.google.com/a/list.nist.gov/g/pqc-forum/c/4iaf...
[4] - https://ntruprime.cr.yp.to/index.html
[5] - https://groups.google.com/a/list.nist.gov/g/pqc-forum/c/7Whv...
[6] - https://groups.google.com/a/list.nist.gov/g/pqc-forum/c/KFgw...
[7] - https://groups.google.com/a/list.nist.gov/g/pqc-forum/c/Fm4c...
[8] - [PDF] - https://csrc.nist.gov/csrc/media/Projects/post-quantum-crypt...
Fun fact, that is because they ARE primarily cryptography people! Goldwasser is known for Blum-Goldwasser or Goldwasser-Micali crypto systems, while Vaikuntanathan is known for Zero Knowledge computations, both materials from any standard cryptography textbook!
(And they're great teachers, I was lucky enough to have them both as teachers in a class a few years back :) )
https://youtu.be/ajGX7odA87k "Why do keynote speakers keep suggesting that improving security is possible?"
If anything, the neural network is more debuggable since you can verify that the decision process you're analyzing (even if complex and hard to understand) was the one actually used in this decision and the same as used for all the other decisions
We already know how frustrating, depressing and dehumanising it can be to experience corporate negligence, where responsibility is diffused to such an extent that it becomes meaningless.
AI will magnify this frustration a thousand-fold unless we acknowledge this problem and put the brakes on AI deployment until we work out how to fix it. And it may be that the problem is insoluble.
IBM research has been looking at data model poisoning for some time and open sourced an Adversarial Robustness Toolbox [0]. They also made a game to find a backdoor [1]
has anyone done this or anything like it?
On the other hand, a seemingly benign perturbation does not necessarily correlate with it being imperceptible. A larger, visually obvious perturbation with a plausible explanation can be less suspicious than a smaller but weirder perturbation.
i suppose the real goal would be a training procedure that tries to ignore stuff outside of the human percept. metamers, masking, noise and attention... oh my.
You should really consider things from a “what can humans perceive” standpoint. There are things you can do with ML and eye saccades that you will literally never see because of perceptual delay. If I can push a saccadic event below 50ms you will never notice it. https://en.wikipedia.org/wiki/Saccade
That’s one example.
"Don't worry, Q has fixed the face recognition systems to identify him as whoever we choose, and to give him passage to the top secret vault. But it would help if if he would just shut up for a while".
That's where the real money is at: subtle AI bot armies that remain invisible yet influence other more public AI systems in ways that can never be discovered. This is the kind of thing that if you ever hear about it, it's failed.
We're entering a new world in which computation is predicable but computational models are not. That's going to require new ways of reasoning about behavior at scale.
One might suggest that the term 'model' is in fact an extremely bad choice of name for the concept of a collection of condensed post-training decision support data in the machine learning world, because it implies a faux-scientific air of objectivity, precision, peer review, and intelligibility for inspection that is entirely undue. IMHO better terminology would have been a new/clean term without conceptual baggage that included some recognition of its fundamental nature: computed/derived/one-way/known-fallible.
There are only two hard things in Computer Science: off by one errors, cache invalidation and naming things. - Phil Karlton
Quotes via https://github.com/globalcitizen/taoup
With companies like Atlassian just going down and not coming, one wonders whether the concept of a technical Ponzi Scheme and technical collapse might be the next thing after technical and it seems like the fragile ML would more accelerate than stop such a scenario.
Edit: Now that I think about it, can't data poisoning happen when predicting, rather than just happening in the training phase? In that case, it's going to be complicated to work around that.