Images that fool computer vision raise security concerns
news.cornell.edu
news.cornell.edu
It reminds me of the CV dazzle anti-facial-recognition makeup that made the rounds a while ago: http://www.theatlantic.com/features/archive/2014/07/makeup/3...
This definitely reinforces my belief that having humans in the loop is not only desirable but necessary. For the majority of human history minus a few years you could only be accused of a crime by another human being. I'd like to see the trend of automated "enforcement" reversed and codify into law that you MUST be accused by a human being.
If everyone is breaking so many laws that the police and courts can't keep up it doesn't mean that humanity is broken. It means that the law has gotten so far out of sync with humanity that the law is broken. People make the laws, not the other way around.
The world would be a much better place if more people realized this.
If the system doesn't have the resources to give every accused criminal a fair trial, then either you're making too many criminals, the system doesn't have enough resources, or both. Bypassing trials is just a way to cover your ears and shout "la la la" to ignore the problem.
Yes, yes there is. We start with the most victimless crimes (drug possession for personal consumption) and go from there.
Police unions will fight it. Prison unions will fight it. The court system will fight it. Nevertheless, it must be done.
Did I not call that out enough? My apologies if so. Stop wasting time prosecuting what doesn't provide value in prosecuting.
There's really no reason why diffuse harms should not be criminal while focused ones should (arguably, diffuse harms are less able to be addressed by individual action through the civil justice system, so outside of focused harms where the victim is unable to pursue action after the crime, like murder, diffuse harms are the most important for government to address directly.)
OTOH, there's often a lot more disagreement over whether what some people perceive as a diffuse harm actually is a harm.
For example, dumping mercury into a major river causes extremely diffuse harm, but I'd never describe it as a "victimless crime."
In the other direction, smoking weed in private harms nobody except the smoker. The harm is so focused it doesn't even touch anyone besides the offender. Yet this is almost the canonical example of "victimless crime."
It's not about diffuse harm, it's about whether anyone was harmed at all besides the accused.
This is precisely the point that is contention. The entire argument for prohibiting marijuana is that this is, in fact, not the case, and that, through a number of indirect channels, people "smoking weed in private" harms others throughout society in a variety of ways.
Obviously, as I said, there is considerable disagreement about whether this is, in fact, the case, (and, also, as to whether, even if it is, prohibition mitigates or exacerbates these harms, and as to whether, in any case, prohibition is an ethically-acceptable response even if the harms exist and prohibition mitigates them.)
But it certainly is not the case that those who support prohibition generally agree that the crimes are "victimless" that those opposed to prohibition describe that way.
Lots of what I would call "victimless crimes" are outlawed because of what the proponents of criminalization see as focused harm. Prostitution, for example, is seen as either harming the prostitute, or harming the patron's family. For drugs, it's often considered that the harm is focused on the user, not diffuse.
And again, lots of crimes with diffuse harm are generally agreed on to not fall into the "victimless" category, like dumping toxic materials.
There is some truth to this for some drugs, but not even close to most.
My point is just that there is no connection, as far as I can see, between crimes which some people as presenting diffuse harm, and crimes which other people see as being victimless. All combinations are not only present but common.
>In the other direction, smoking weed in private harms nobody except the smoker. The harm is so focused it doesn't even touch anyone besides the offender. Yet this is almost the canonical example of "victimless crime."
It's true that this doesn't directly harm anyone. However, if the smoker is doing this 'illegally' (without a medical license, not in WA or CO, etc) and didn't grow it, he/she is participating in and supporting an illegal drug market via increased demand. If the smoker weren't participating, it would reduce the demand that drives smugglers and the horrific things in Mexico.
It is rather looking at the overall picture of "What are the consequences of allowing weed to be bought/sold/grown. Is there harm in it?" The whole infrastructure, not just the end user.
Increasing resources for holding trials is also not terribly difficult and doesn't require any sort of overhaul. It's just a problem of money, and not a big one. Just for a random example, it looks like the courts account for about 0.3% of my local county budget and about 1% of my local state budget. We could literally increase court resources by a factor of 10 with only a modest increase in taxes to fund it.
It might be that there's something inefficient in the trial process that holds no weight on "fairness". I'm not a legal expert by any means, so it may not be the case, but it seems as though it's a possibility.
On the other hand, if you can get the number of accused criminals down to a reasonable level, it wouldn't matter too much if there was waste in the smaller number of trials.
No doubt the problem can be attacked from many directions.
How do we address this?
One emerging solution employs advanced AIs to pore over the reams of evidence in order to better inform lawyers of the legal situation. I can imagine extrapolating this particular methodology far enough into the future such that the entire process subsequent to arrest (processing, pre-trial hearing, trial and sentencing) might be automated to the degree that it could be accomplished within minutes instead of months.
What would society look like if we had such an "efficient" judicial system?
That's a potentially terrifying possibility. It may be that we'd end up incarcerating orders of magnitude more people; a dystopian reality. How would we then rein in such a powerful system? I have no idea.
Not just that, though -- it also comes from the inherent inefficiencies in trying to recover exactly what happened from any given situation in which the question of whether a crime was committed.
We could start recording everything that happens, but... that's also a potentially terrifying possibility.
You're in luck! Facebook, Google, and many of the wonderful, selfless people who post on HN are already working on that... or at least defending others' "right" to do so if they aren't doing it themselves.
Various entities are tracking who you know, who you sleep with (http://www.whosdrivingyou.org/blog/ubers-deleted-rides-of-gl...), what your face looks like, what your friends' faces look like (thanks, photo tagging enthusiasts!), who you talk to, what you say to them, when you say it, where you go (thanks to various wonderful sources, including ALPR companies like Vigilant), how long you stay there, what you eat, what you wear, what you watch, what you listen to, where you move your mouse while viewing websites, who your doctor is, what medications you take, what the symptoms of that last rash you had were, what your political views are, what you read, where you work, how many steps you took today, what websites you visit, and about 30,000 other bits of data... just to keep you safe!
The future is so amazing! I don't know what I would do without a customized advertising experience(tm). Such a drastic improvement over life in the past where people were so bored of advertisements that they chose to avoid watching them! What none of us knew at the time was that we really just wanted to see more relevant advertisements more often, while giving up our privacy for the corporations' greater good! I sleep much more soundly after a solid day of being bombarded with advertisements that teach me to be a better consumer!
80% of property crime goes unsolved, so unless you legalize vandalism and theft there are going to be criminals.
But wait. He specifically exempted police unions from his busting legislation, didn't he?
http://www.washingtonpost.com/blogs/govbeat/wp/2015/02/18/sc...
> To help close the state’s $283 million budget shortfall this year, Wisconsin Governor Scott Walker (R) plans to skip a $108 million debt payment scheduled for May.
> By missing the May payment, Walker will incur about $1.1 million in additional interest fees between 2015 and 2017. The $108 million debt will continue to live on the books; Walker’s budget proposal for 2015-2017 will pay down no more than about $18 million of the principal.
> In March last year, Walker signed a $541 million tax cut for both families and businesses. At that point, Wisconsin was facing a $1 billion budget surplus through June 2015, the Journal Sentinel reported.
The elected official's opponent in the following election simply runs a "JOHN SMITH IS SOFT ON CRIME, JOHN SMITH IS BAD FOR US AND OUR FAMILIES" campaign, and people who think public-sector unions are the devil will still vote for the opponent, because now you're pressing their bias (which is in favor of "tough on crime, lock 'em all up and throw away the key").
It's basically the same as setting up your conditions to take advantage of short-circuit evaluation. You don't put the time-consuming and resource-hungry part first.
Were there even the slimmest chance of acquittal, few defendants would utter that phrase without there being either a benefit to owning up, or an extra penalty for not doing so. This is what plea bargaining and TICs are for.
If there's a suspicion of false confession, meaning that the suspect may not be the offender, that should normally be sorted out before court, so that the correct offender is tried for the correct offence (e.g. wasting police time if voluntary, or some kind of intimidation/coercion offence if not)?
How? Court is the process for sorting that out and assessing whether there is doubt, and the extent of that doubt.
> If there's a suspicion of false confession, meaning that the suspect may not be the offender
Are you assuming that this represents a minority of cases? Who has to suspect that the accused is not really guilty, a jury of his peers?
In reality it becomes a punishment for demanding a fair trial.
Oh man, that's so wrong. Even where there is a LARGE change of acquittal, many choose a plea bargain because they cannot afford a good attorney or because the prosecutor is threatening them with something crazy like 40 years for downloading a movie. Would you take the risk of 40 years, knowing you're innocent? Do you have enough faith in a jury to put the rest of your life in their hands? I doubt it.
That said, I don't believe it's very difficult to make an ethical and legal case against the concept of plea bargaining in the first place. I would argue it's very important to a free society to waste everyone's time proving that an accused person is guilty, and that not doing so violates the sixth amendment and fundamental human rights.
But we should not reward them for doing so. Never should a person be presented with a choice between a certain lesser punishment, or a fair trail and a potential greater punishment. That's the "bargain" in "plea bargain," and it's completely reprehensible.
You say, "few defendants would utter that phrase without there being either a benefit to owning up, or an extra penalty for not doing so." That's exactly how it is! A plea bargain isn't just, "we both know you did it, so just confess and let's skip all the lawyers and stuff." It's always, "We both know you did it, so confess and we'll let you out early. If you insist on taking this to trial then we will throw the book at you." A lot of innocent people will take the plea when faced with that choice.
Not every crime needs to go to trial, but every accused criminal needs to have the right to a trial without being punished for exercising that right. If you remove the punishment then you'll no longer have plea bargains, just pleas.
I think that a partial measure towards reforming plea bargains would be to require the police to present their evidence for the court to review before entry of a plea. This creates a public record that somebody could investigate in the future. There could also be a provision that if exculptatory evidence is revealed in the future, the plea can be rescinded.
The other issue is police overcharging crimes. For example, suppose you, in a fit of pique, cut your roommate's arm with a knife. Not cool; that's aggravated assault. But the police then charges you with attempted second degree murder -- or even first degree (eg, premeditated). Now your defense attorney has to work hard just to get your charges down to a reasonable level; the agg assault you should have been charged with in the first place. This is what a huge majority of plea bargaining is about; not getting away with it, but getting the charge down to something that describes your actual offense.
(Source: longtime criminal defense nerd)
By the 70s crime had skyrocketed and it was clear that they had gone too far. But instead of issuing a mea culpa and reexamining past rulings, the various courts started allowing prosecutors to claim broad new powers and take extremely aggressive tactics.
By this point there's no real way to fix it. The legal system is based strongly on precedent. It can't just undo major rulings of the past and replace the with something sane.
sounds exactly like the point where a major fix must be done. Life changes faster and faster. The precedent system was great in 15th century when things look pretty much the same like in the 14th or 13th century.
Any claim that court ruler enabled crime is highly debatable. Factors causing increases and decreases in crime rates are difficult to pin-point.
One argument is returning Vietnam veterans lead to the increase in violent crime - post-war eras have a history of being associated with crime-increases from the return of veterans habituated to violence.
Um, so? Then we need to allocate more resources to prosecute cases.
If I am falsely accused, I want my day in court. And I want it to be fair. The current system has problems on both fronts.
The problem is that there's no way for anyone else to determine a priori if your accusation is false or true. That's what the whole presumed innocent until proven guilty thing is about.
In reality, if you are accused AT ALL, you want your day in court and you want it to be fair. Even if you had committed a crime, if the police did something they're not supposed to that needs to get sussed out in court and you should go free.
Half the point of a trial is to make sure that nothing unfair is done by the investigators (police, prosecutor, etc). This is to keep their power in check so that they'll follow the rules. Otherwise it could get mighty tempting to fudge something a little bit "because we KNOW this is the guy!" and "we need to do the right thing."
> The legal system is based strongly on precedent. It can't just undo major rulings of the past and replace the with something sane.
Sure it can; major court rulings have been overturned by later courts, and, even if courts won't do that, the law which they were interpreting when they made the rulings (including the Constitution) can be changed, rendering the old rulings moot.
But that's just the anarchist in me speaking.
They just forget it as soon as the news media present whatever is the current horror story.
in other words, the bandwidth of the image recognition product is not the same as the bandwidth of the video storage.
> the "speed" dial would be turned up as high as necessary
It's possible, and probable given the situation (i.e. government, terror acts), but are there situations where this wouldn't happen in a different circumstance? Would mall security, or a marketing company operating an advertising product, have the ability, authorization, know-how, or financial incentive to do so if this wasn't a life/death situation?
Not every application of facial recognition technology will be targeted to terror suspects/bombings.
"They tested with two widely used DNN systems that have been trained on massive image databases."
Is this wrong?
I had a similar thought. An app which takes the image of a target - Julian Assange say - and overlays it onto your facial image as thousands of virtual stickers. Gradually the evolutionary algorithm morphs this overlay towards your real facial image; its fitness function being fewer, more unobtrusive stickers. All the while maintaining compatibility with target facial recognition.
One day, J.A. escapes from his embassy lair. Thousands of people simultaneously appear on the streets wearing physical stickers on their faces, addressing the London panopticon: "I am SpartAssange".
In particular, this weakness has nothing to do with Computer Vision and also nothing to do with deep learning. They only break ConvNets on images because images are fun to look at and ConvNets are state of the art. But at its core, the weakness is related to use of linear functions. In fact, you can break a simple linear classifier (e.g. Softmax Classifier or Logistic Regression) in just the same way. And you could similarly break speech recognition systems, etc. I covered this in CS231n in "Visualizing/Understanding ConvNets" lecture, slides around #50 (http://vision.stanford.edu/teaching/cs231n/slides/lecture8.p...).
The way I like to think about this is that for any input (e.g. an image), imagine there are billion tiny noise patterns you could add to the input. The vast majority in hundreds of billions are harmless and don't change the classifications, but given the weights of the network, backpropagation allows us to efficiently compute (with dynamic programming, basically) exactly the single most damaging noise pattern out of all billions.
All that being said, this is a concern and people are working on fixing it.
Agreed; the weaknesses reported should definitely not be taken to affect only convnets or only deep learning. Ian's "Explaining and Harnessing Adversarial Examples" paper (linked by @Houshalter) should be required reading :).
> backpropagation allows us to efficiently compute (with dynamic programming, basically) exactly the single most damaging noise pattern out of all billions.
True. By using backprop, one can easily compute exact patterns of pixelwise noise to add to an image to produce arbitrary desired output changes. However, it's an important detail that that most of the images in the paper (all except the last section) were produced without knowledge of the weights of the network or by using backpropagation at all. This means a would-be-adversary need not have access to the complete model, only a method of running many examples through the network and checking the outputs.
> ...there are billion tiny noise patterns you could add to the input.
Perhaps because the CPPN fooling images were created in a different way (without using backprop), they seem to fool networks in a more robust way than one might think. Far from being a brittle addition of a very precise, pixelwise noise pattern, many fooling images are robust enough that their classification holds up even under rather severe distortions, such as using a cell phone camera to take a photo of the pdf displayed on a monitor and then running it through an AlexNet trained with a different random seed (photo cred: Dileep George):
http://s.yosinski.com/jetpac_digitalclock.jpg http://s.yosinski.com/jetpac_greensnake.jpg http://s.yosinski.com/jetpac_stethoscope.jpg
I thought this was surprising the first time I saw it.
Of course, if you're as clever as Tinbergen, you might be able to come up with patterns that fool organisms even without (1) or (2):
A sort of related idea was explored with this site, where millions (ok, thousands) of users evolve shapes that look like whatever they want, but likely with a strong bias toward shapes recognizable to humans. Over time, many common motifs arise:
You'd need another classifier to tell you "nope it's actually just random noise and shapes" ... hm.
Who gets to decide what is really a picture of a panda?
If we'd manage to craft a picture that could with very high certainty trick human neural nets (for the sake of argument, including those higher cognitive functions) into believing something is a picture of a panda, "except it actually really isn't", what does that even mean?
Human insists it's a picture of a panda, computer classifier maintains it's noise and shapes.
Who is right? :)
Kind of like the memes in Snowcrash, ancient forgotten symbols that make up the kernel of human thought.
That would require that human vision works in the fashion of neural networks/svms/etc. There actually isn't any evidence for this.
Likewise, I was not very surprised that you can produce fooling images, but it is surprising and concerning that they generalize across models. It seems that there are entire, huge fooling subspaces of the input space, not just fooling images as points. And that these subspaces overlap a lot from one net to another, likely since they share similar training data (?) unclear. Anyway, really cool work :)
Agreed. That is surprising, and also increases the security risks, because I can produce images on my in-house network and then take them out into the world to fool other networks without even having access to the outputs of those networks.
> likely since they share similar training data (?) unclear.
The original Szegedy et al. paper shows that these sort of examples generalize even to networks trained on different subsets of the data (and with different architectures).
> Anyway, really cool work :)
Thanks. :-)
Good point. You could also do this with the gradient version too (fool in-house using gradients -> hopefully fool someone else's network), but the transferability of fooling examples might differ depending on how they are found.
But I'm having problems generating images: A top SVM classifier on MNIST has a very stable confidence distribution against noisy images. If I generate 1000 random images, only 1 or 2 of them will have confidences that are different from the median confidence distribution. That is, all the images are classified with the same confidence as class 1. They also share the same confidence for class 2, etc.
So it is very difficult to make changes that affect the output of the classifier.
Any tips on how to get started with generating the images?
Yep, true, might just take a while. On the other hand, even a very noisy estimate of the gradient might suffice, which could be faster to obtain. Perhaps someone will do that experiment soon. Maybe you could convince one of those students of yours to do this for extra credit?? ;).
> Likewise, I was not very surprised that you can produce fooling images, but it is surprising and concerning that they generalize across models.
Ditto x2.
> It seems that there are entire, huge fooling subspaces of the input space, not just fooling images as points. And that these subspaces overlap a lot from one net to another, likely since they share similar training data (?) unclear.
Yeah. I wonder if the subspaces found using non-gradient based exploration end up being either larger or overlapping more between networks than those found (more easily) with the gradient. Would be another interesting followup experiment.
I agree with most of what you say, but note that nearly all of the images in the paper were generated without the gradient. I.e. all the images produced by evolution did not use the gradient, only the output of the network regarding its prediction confidence. There are some images that use the gradient, but only to show a 3rd class of "fooling images".
PS. It's nice to see our work (both this paper and the NIPS paper on transfer learning) in your class. Thanks for including it. I wish I could have my students take your course!
Basically neural networks and many other machine learning methods are highly linear and continuous. So changing an input just slightly should change the output just slightly. If you change all of the inputs slightly in just the right directions, you can manipulate the output arbitrarily.
These images are highly optimized for this effect and unlikely to occur by random chance. Adding random noise to images doesn't seem to cause it, because for every pixel changed in the right direction, another is changed in the wrong direction.
The researchers found a quick method of generating these images, and found that training on them improved the net a lot. Not just on the adversarial examples.
No they're not! You introduce non-linearities like the sigmoid or tanh to make them highly non-linear.
>The linear view of adversarial examples suggests a fast way of generating them. We hypothesize that neural networks are too linear to resist linear adversarial perturbation. LSTMs (Hochreiter & Schmidhuber, 1997), ReLUs (Jarrett et al., 2009; Glorot et al., 2011), and maxout networks (Goodfellow et al., 2013c) are all intentionally designed to behave in very linear ways, so that they are easier to optimize. More nonlinear models such as sigmoid networks are carefully tuned to spend most of their time in the non-saturating, more linear regime for the same reason. This linear behavior suggests that cheap, analytical perturbations of a linear model should also damage neural networks.
If x is a bit-vector then this can be as simple as saying "flip one bit of the input and here's how to predict which output bits get flipped." When you're building a hash function in cryptography, you try to push the algorithm towards a non-answer here: about half the bits should get flipped, and you shouldn't be able to predict which they are. But of course there's a security vulnerability even if + and ⊕ are not XORs.
Resisting "adversarial perturbation" in this context means basically that neural nets need to behave a bit more like hash functions, otherwise they will confuse the heck out of us. The problem is that if you just took the core lesson of hash functions -- create some sort of "round function" `r` so that the result is r(r(r(...r(x, 1)..., n - 2), n - 1), n) -- seems like it'd be really hard to invent learning algorithms to tune.
But it doesn't really matter what activation function you use. The paper argues that its the linear layers between the nonlinearities that are the problem.
In every security system, it is assumed that if there is an attack surface, sooner or later an intelligent adversary will come and exploit it. And there is a long precedent saying that if you see the word "linear" anywhere in the attack surface description, the adversary is bond to come sooner rather than later.
http://www.evolvingai.org/fooling
Some of them are very simple, and DO occur a lot in the world. For example, the alternating yellow and black line pattern would be encountered by a driverless car, and it would think it is seeing a school bus.
While the image shows a yellow and black line pattern to us, are you sure this is also what the CNN "sees"? Couldn't this image just be the same as the adversarial images, i.e. it responds to many small input values rather than the overall pattern?
If it's possible to make the CNN predict an ostrich for an image of a car, then the same can be done of an image of an alternating yellow and black line pattern, no?
Would a preferred computer vision system experience the Checker shadow illusion? http://en.wikipedia.org/wiki/Checker_shadow_illusion
If yes, computer vision will be as fallible as ours. If no, then there will always be examples, like the ones presented, where computers will see something different than humans.
People recognize things by building a 3D model in their head, then comparing that to billions of experiential models, finding a match and then using cognition to test that match. "Is that a bird? No, its just a pattern of dog dropping smeared on a bench. Ha ha!"
I meant to talk about what some hypothetical future system could do (which I think was a reasonable context given the comment I replied to), not to characterize current systems.
To get there, computers will clearly have to change utterly their approach. A cascaded approach of quick-math followed by a more 'cognitive' approach on possible matches, could definitely improve on the current state of affairs.
I can't help but believe some of the image recognition mentioned in your article, especially of icons, is built through previous experience with similar iconic images. Symbols for things become associated with the real things. Its a modern adaptation of a much older processing mechanism.
So you're saying people are generative reasoners with very general hypothesis classes rather than discriminative learners with tiny hypothesis classes.
To which the obvious response is, yes, we know that. The question is how to make some computerization of general, generative learning work fast and well.
That being said, here is a much better resource than I am: http://en.wikipedia.org/wiki/G%C3%B6del%27s_incompleteness_t...
When we look at the Checker Shadow Illusion, our brains are automatically "parsing" that image into a 3D scene and compensating for the lighting and shadowing. The reason you see square B as being lighter than A is because, if you could reach into the image and remove the cylinder so it's not casting a shadow anymore, square B would be lighter than A. Our brain doesn't think about color in terms of absolute hex values. Instead it tries to compensate for the lighting and positioning of elements, assuming that they're similar to what we would see in real life—a challenge which, by the way, computer vision systems have always struggled with.
That sounds pretty "correct" to me.
I don't know the answer. I do find it fascinating that something as simple as perspective (which we take for granted in a graphic like the checkers illusion) is a fairly recent technique, invented by Renaissance artists.
Furthermore, I am seeing the security concerns, but I figure this is far from a practical attack. Deep Learning Classifiers do not act as gatekeepers: You have not much to gain from a single faulty classification. You won't be granted access to secret information if you happen to look like the CEO.
Perhaps it's a sign of the times that almost every discovery that could possibly be related to security in some way, does. I have a feeling that if this was a decade or two ago, the sentiment would be very different. ("Can you figure out what a computer thinks these images are?")
Also, the image labeled "baseball" immediately reminded me of a baseball...
Sure, but this assumes that whatever neural net system you're relying on was bought by someone who is more security conscious than they are cheap.
Trivially "solved" by treating the ensemble as a single object, then constructing a counterexample. My intuition suggests that while the resulting "fooled you" image may very slowly converge on something human recoginizable, it won't do so at a computationally-useful rate.
If your intuition is right though, then the ensemble may be able to counter with a random selection of nets for its vote: You'd need to evolve images for every possible combination and/or account for nets added in the future.
I believe the original paper tried ensembles and even got the images to work on different networks.
The original paper asked if images that could fool DBN.a could fool DBN.b. The answer was: certainly not all the time. They used the exact same train set and architecture for DBN.a and DBN.b, just randomly varied initial weights. I think this is too favorable for a comparison with a voting ensemble made with nets with a different architecture, train set and tuning. Can they also find images that can fool DBN.a-z?
Also, to test if a net can learn to recognize these fooling images, they simply add them to the train sets. Those noisy images would be far simpler to detect: They have a much greater complexity than natural images. To detect the artsy images, a quick knearest-neighbors run should show that they do not look much like anything it has seen before, so it may be an adversarial image.
>In addition, the specific nature of these perturbations is not a random artifact of learning: the same perturbation can cause a different network, that was trained on a different subset of the dataset, to misclassify the same input.
a relatively large fraction of examples will be misclassified by networks trained from scratch with different hyper-parameters (number of layers, regularization or initial weights). The above observations suggest that adversarial examples are somewhat universal...
Tell that to this startup: http://onevisage.com/ovi-technology-page/
Both results are interesting from the point of DNN construction, and there have been some papers suggesting ways to counter the effects specified in the Svegedy research. In practice (as others have mentioned) in order to construct an exploit similar to the one described in this post, you'd need to have a lot of knowledge about the DNN (e.g. weights) that an external attacker wouldn't have.
What this does leave open, though is a disturbing way for someone with internal access to a DNN doing important work (e.g. object recognition in a self-driving car) to cause significant damage.
Now seriously, modern cameras have face recognition and ton of other features. It should be possible to cause crash.
https://en.wikipedia.org/wiki/EURion_constellation#Other_ban...
In particular, have a look at this:
Digital camera sensors output a stream of bits that depend on the intensity of the light reaching the pixels. For normal images, there is (relatively speaking) not so much contrast, so the signal has few high-frequency or repetitive components to it. However, if you point the sensor at an image that effectively causes each pixel to be alternating from full dark to full bright, the signal becomes far more regular and the high-frequency components increase significantly. A possible problem is that, since the bulk of the power draw occurs when a bit transitions from 0-1/1-0, this repetitive and high-frequency signal causes more stress on the power supply circuitry (look up "voltage regulator oscillation"), and if it causes voltages to go out of tolerance, can crash the system. I suppose an image that produced a stream of 010101010... for each pixel's value could also be an example of this. Other resonant effects may also play a role in this; if the oscillation frequency happens to synchronise with something else, physical damage is a possibility if the components are pushed beyond absolute maximum ratings. It's an extreme edge case, not normally encountered in use.
The reason why I think this could be plausible is that, although I've never tried/experienced this with a digital camera, I had an old analogue video camera that would work fine in all circumstances except when pointed at a monitor displaying its image, upon which it would emit a loud high-pitched whine and then shut itself off. I discovered that one of the power supply rails would go into oscillation when the camera saw itslf, and this was enough to shutdown the system. Adding some extra supply decoupling was enough to stop this from happening, but apparently this problem has been known for a while:
Which is not to say that we should all fear computers more than humans as a consequence; we do inexplicable things, too.
This seems markedly different from biological neural networks. Is the difference one of network structure/algorithm or rather the fact that biological neural networks (in human image processing) actually have time and space to learn a lot about each individual image class?
It is, however, possible to have a deep net produce 3D models/images: https://www.youtube.com/watch?v=QCSW4isBDL0 "Learning to Generate Chairs with Convolutional Neural Networks".
I also suspect a different part of cognition is used when humans are asked to recreate a "fire truck" than when humans are asked to classify a "fire truck" from a "car". The former seems closer to using memory ("what did the last five fire trucks I saw look like?"). A fairly recent addition to deep nets is making use of memory: http://arxiv.org/pdf/1410.5401.pdf "Neural Turing Machines". So the difference may quickly become less significant.
But I'm not clear how important this phenomenon really is to the practice of CV, since 1) 'spoofed' images are highly specific to each DNN being used, and 2) a trivial reality check of the image can always 'out' examples like these.
Also the question is raised as to whether or not new methods of spoofing are possible that aren't so easily detectable.
This means that come the singularity AIs will have to use AI specific CAPTCHAs in order to distinguish between humans (aided by dumb computers) and other AIs.
A C, however, can be drawn as half of a circle - one command to draw an arc, a center point, a radius, and the start and stop angles. Five pieces of data.
(This is assuming, of course, that computers prefer minimal amounts of data. If that's wrong, then the computer would obviously prefer B. You need more data to describe it.)
Also, one could make CAPTCHA's incredibly hard. Humans and dumb computers won't be able to solve it, and singular AI's get a pass. So give access to anyone who isn't able to solve the CAPTCHA and redirect the singularity AIs to google.com :).
Finally, the singular AI could create a CAPTCHA which separates humans from AI. The problem of separating humans from AI will quickly become too difficult for humans to solve.
I get a mental image of a future where computers try to keep us pesky humans out of their websites.
I've often thought there would be an awesome opportunity in there to make a hilarious app that catches cheaters during the music round of Pub Quiz.
The best pub quizzes dim the lights so that smartphone cheats beam forth their sneakiness.
The image recognition in Brain Age is obviously much simpler, but it's still basically the same idea.
http://yosinski.com/media/papers/Nguyen__2014__arXiv__Deep_N...
As long as you don't publish the specs of your net you should be fine I guess.
An oracle, here, would be any version of the system that you can query against repeatedly without suffering too much of a penalty.
The purpose of this research is to work as a proof of concept. Sure, this iteration needs to cheat slightly to achieve its results. However, looking at the images that aren't just noise, it seems to be possible to construct less specialized images that fool many nets.
No because the instant you used an image that was close, but wrong the human brain would retrain.
I wonder in the neural net can do the same - try this experiment on an active neural net and let it train itself. (i.e. don't tell it it's being faked, let it figure it out then correct for it).
On the other hand, a computer mechanically relates the specific format of a keyboard to the word "keyboard." It fuzzy matches the pixels of images to extract the object in the image: not the individual ideas implicit in the image.
Computer Vision needs more depth to actually be considered vision.
We use ML-based computer vision at my work, so I have a bit of experience here. I think the biggest practice take away from observation the ML can give some wonky results is that ML system can be a real PITA to debug.
On a serious not couldn't we keep training the same DNN's using these white noise images as negative examples?
>In a further step, the researchers tried ?retraining? the DNN by showing it fooling images and labeling them as such. This produced some improvement, but the researchers said that even these new, retrained networks often could be fooled.
Could autistic children have learning impairment due to their inability to correctly sort/segregate stimuli, in the same way that these neural networks generate high-confidence false-positives?
The algorithm would be useful if it could generate/identify images that feel as emotionally transformative as good art.
There is also a video summary of the paper there.