Show HN: I am a high school student and I did research on 3D adversarial attacks
arxiv.org
arxiv.org
A high-level overview of the research:
Basically neural networks are weak against adversarial attacks that change the input by a little bit to cause the prediction to be wrong. We look at these adversarial attacks in 3D space, specifically on 3D point clouds (think LiDAR and RGB-D data). In the paper, four attacks in two different categories (distributional and shape attacks) are proposed. The main benefit of distributional attacks is their imperceptibility. On the other hand, shape attacks are more easily crafted in real-life (though more perceptible) and also robust against point removal defenses that were proposed in previous work. If you want a more comprehensive (but less dense than the paper) overview, take a look at my blog post [2].
Joking aside, fantastic work. The differentiable reformulation in the distributional attack is a tour de force.
A question. Why did you do it?
I do research because I like solving hard problems that people have never considered. I like to ensure that what I have learned will be put into practical use.
Why would 3d not be a subcase of that work?
Similarly, adversarial attacks/defenses are still being proposed for graphs, audio, and other domains because we can leverage domain-specific knowledge.
Would you have a canonical name for this distribution ? If you try matching log likelihoods, what parametric family does it resemble ? Briefly, given one of the canonical two dozen (uni/multi)variate distribution, one can create new distributions either by location-scale transform, mixtures, or say by using a k-param EFD family. So if I pick a k-param MVN ( multivariate normal with k means, k sigmas & O(k^2) correlations, I can create new distributions all day long by tweaking these 2k+k^2 params until cows come home. Brittle inference engines such as CNNs trained on a specific family with specific (hyper)parameters will fail once the distribution changes significantly, though visually the changes will be imperceptible.
In my previous paper, I've shown that moving the points around on the surface of an object does lead to imperceptible but effective adversarial attacks, as you've observed.
I'm unconvinced by this statement. There are many attempts to negate attacks that do so by applying linear transformations, masks, etc. To images. Removing pixels is not novel.
We like to imply that domain knowledge is relevant but after you design a feature vector it all ends up the same.
The time dimension adds complexity to the problem as the optimal values for the perturbation vary depending on both the immediately surrounding values, and many of the values beforehand.
When I say “hello world”, the fact I said “e” depends on the fact I said “h”. “L” depends on both “e” and “h”... etc etc.
Adds an extra dimension to the problem.
Also, distance metrics for images aren’t ideal for audio, for many reasons. That’s why audio signal processing is a different sub field vs image processing.
The approaches are similar, but we have to use different things in the end because audio behaves differently to images. Eg feature extraction through MFCC is a variant of Fourier, but specifically tailored for the human ear.
E.g. Lea Schonherr et al.’s really good Psychoacoustic attack paper.
On the negation of attacks through transforms - important to remember that an ensemble of weak defences are not strong. Many attacks have been shown to be robust to simple transformations.
This paper does suggest that we can circumvent certain domain-specific knowledge when attacking. This does not mean that we won't discover methods to utilize domain-specific knowledge in the future. I would imagine extending current provably robust methods to 3D would require domain-specific knowledge to deal with the distribution of points.
Are the attacks statistically identifiable? Can they be translated to a training subset? What would you propose?
I am not sure about statistical identification, but we show that it is difficult to identify and remove adversarial points by looking for statistical outliers points.
I am not sure about truly robust 3D-specific defenses---if anyone has some idea, I am open to collaboration. I would imagine some sort of provably robust method built specifically to handle the varying density and distribution of points.
As a fellow high school student who has a good amount of experience and knowledge in deep learning, how would you recommend me to move forward? I am struggling to find opportunities to show off my work and knowledge, and would like some advice.
Otherwise, try reimplementing algorithms and blogging about them. Do fun projects like deploying a model online or to phones. If you are a fan of competitions, then you can try some Kaggle competitions. With some projects (if you say you have experience then you probably already done some) it should not be hard to get research experiences because you have something to show off. Remember to post your projects on Reddit and HackerNews to get internet points and encouragement! It is quite motivating.
That’s where I (fellow High School student) struggle. I’m not really at a point where I can contribute much either way.
What was that luck?
Here's what I think when he says "luck": I emailed 50+ personalized emails over a span of 2-3 months to professors and anyone else I wanted to work with. Most didn't respond, but for those who did, you'll need some luck and networking skills (which I hope have improved) to convince them you're the right fit or can help without hassle.
Specifically luck includes: whether or not that person had a good day or had enough time to check their email, or numerous other things
I'm not particularly a genius, but deeply passionate and curious on what I do. As a result, doing research or an internship was more fun than most other things I'd be doing over the summer/fall.
Didn't grow up in the best of situations, but learned how lucky I was to live in the Bay, giving me access to meet people IRL and grow from there.
The two are often indistinguishable from each other.
You have a very wise perspective for your age. Good luck to you.
First, fantastic that you're doing academic work so early. I believe most students wait far too long to be exposed to this aspect of academia (and their lives), which is far more about asking good questions, achieving deep understanding, and getting good results, than memorizing some procedure that isn't necessarily useful (but happens often in school). Keep creative and keep working on creating a good toolset (find math tools you find interesting/useful and own them!).
Now on to the paper. I'd reiterate the importance of asking good questions almost above results. For example, the result of adversarial sticks and sinks looks good -- but is it asking the right question? If you think realistically, adversarial attacks can occur in a number of ways.
One of them is that a human classifies a dataset one way while a machine another. In this case you would also want the human not to be able to tell you data is weird or there's something funky going on -- that is clearly the case with sticks (and sinks to a lesser extent). A human could be easily trained to spot them, and generally tell something weird is going on.
Another attack scenario is where you can modify some object, like a picture, but have some restriction on how much you can modify it. For example, you can manipulate only some bits of an image, or only perturb a small part of a real-world object that is under classification (say by putting a sticker on a car and fooling a system into thinking it is a dog, or something). If there is no restriction on your perturbation, this problem would be trivial (just replace the data with intended object data). The justification behind sticks and sinks does not look very well fundamented.
So sticks/sinks do not fare too well in either case, despite looking very good in terms of success vs defenses (although there's a chance they could inspire more practical attacks).
The commentary on Haussdorf distance is relevant here, but only on the first case (fooling human judgement), and it is of course an imperfect proxy (the true metric is human perception) -- another hint that fundamentals (and applications) are important to keep in mind.
Overall the paper seems well written and I specially like the numerous illustrations.
Keep the good work and don't forget to always look for the inspiring, beautiful and impactful, and seeking understanding. With a little of this in mind I have no doubt you can achieve very much. Good luck!
The success rates of the attacks are not really emphasized---we only show that they work. They provide other benefits like robustness against point removal defenses.
The attacks are optimized for different criterion. The sticks attack is supposed to be easy to construct, at the cost of perceptibility. The distributional attack is more geared towards imperceptibility. Indeed, we do bound the perceptibility of the sticks and sinks attacks, just with different metrics. You can even argue that the number of sticks we generate is a measure of perceptibility. Compared to other papers in terms of visual perceptibility, our attacks are not that crazy. Of course, human perception is the true metric, and I think more work must be done on quantifying perceptibility in 3D. This paper is the first step, and I mainly wanted to show that there are factors other than perceptibility that we care about.