Images altered to trick machine vision can influence humans too
deepmind.google
deepmind.google
Perhaps participants are reusing bouba/kiki[1] skills, evaluating whether the image looks organic (rounded) or inorganic (spiky) - and making their choice accordingly.
>perturbed by a seemingly random pattern
ie. it doesn't really seem like a random pattern. It looks more like a match pattern taken from the mid-layers of a visual NN, the pattern of edges and vertexes which "lights up" more frequently when a cat image is presented to the NN. That would explain why an NN (even another one as they all converge on processing edges and vertexes) would mistake that image of vase for cat - in the altered image of vase those "cat" pattern edges and vertexes are present, just at the attenuated 2 levels amplitude, and the first layers of NN (usually converged to Gabors and Gaussian pulling like in the biological visuals cortex) would still detect those edges and will send it to the mid layers where it will "light up" similar "cat" patterns (it would be very illustrative if the authors provided the mixed-in image subjected to edge detection - that should clearly show a mix of original vase image edges and the added "cat" pattern ones). That also explains why people would point to such an image as more cat-like - the first layers and even the first-mid layers in our visual cortex are very similar to what the first and the mid layers of a well-trained NN converge to, and those our layers would similarly detect and "light up" to those mixed-in cat pattern edges and vertexes - we are just better at signal/noise, ie. clearly recognizing the vase and thus suppressing the "noise" of those cat patterns triggered in our visual cortex, so we wouldn't misclassify the image as a cat, yet we do feel something cat-like. Would be interesting to give that test to the people under hallucinogens, i.e. when the visual cortex mid-layers signals are subjected to much less frontal processing and the "noise" rises to the level of "signal".
IMHO this is not the same as computer vision thinking a rolled over school bus is a snow plow.
This is asking if someone sees an elephant or a unicorn in a cloud.
Asking if a picture of a stop light at an intersection is "cat like" seems to be pretty suspectable to over fitting.
Rorschach inkblot test is pretty much pseudoscience, can someone please explain how this is not similar?
Consider the difference between a human stating "that's an impossible nonsense picture, but if had to describe it then it's a half-Cat and half-Truck abomination" compared to a computer yielding " There is a 50% chance that is a Truck, and a 50% chance that is a Cat."
which is, at least to our conscious verbalizing mind, ~indistinguishable from the left hand picture. the interesting part is they _are_ distinguishable.
as you point out, you'd expect a human to just be like "uhhhh...either?", but it turns out we do see something subconciously, because people do identify the overlaid image at > chance
To me it seems like a conscious "if I had to pick the more cat-like one, this picture has a bit right here that sort of looks cat-like".
i.e. it's weird to say the images "influence" us. It seems more accurate to say that when prompted, we can perceive/understand what is confusing the model.
To me that seems a reasonable--and not very exciting--explanation for why humans may get similar this-or-that answers to a machine-vision model.
A picture of a cat still retains an obvious cat-ness to humans, even when it's been tainted by "truck-ness", but they're slightly more likely to agree it's more truck-y. That response scales with how perturbed it is. (Fig 3 in paper.) If you added the perturbations for "vehicle-ness", would that response be stronger while affecting the image less than cracking up the intensity on the effect? Could you start combining separate concepts, and pick them out individually as like... "activation scores" or something?
If so, feels like compression all of a sudden. I know there's a ton of other ML compression things out there, but that just feels like it could be really information dense.
https://link.springer.com/article/10.3758/BF03206939
Edit: I believe this linked survey is not the subject of the OP.
Instead they barely grasp at straws and come to an obviously inflated conclusion that neural nets and human brains are similar in some way because of this (which _certainly_ isn't something they arrive at or even experiment for, in my opinion).
Aren’t they? Are you sure you understand what is the finding?
> they would be interesting enough if they had managed to collect a large enough sample
They did. The grand parent comment failed to read the right paper.
> obviously inflated conclusion that neural nets and human brains are similar in some way
But that is what they find. The human looks at two almost identical looking images of flowers. And yet when they are asked which one is more cat like they pick the one which the neural network thinks is cat-like too.(Or at least they pick it more often than if they were just selecting randomly in this seemingly nonsense task.) That is exactly “similar in some way”. Similar in which image they find more cat-like. That is the similarity.
Yeah, I got it.
> But that is what they find. The human looks at two almost identical looking images of flowers. And yet when they are asked which one is more cat like they pick the one which the neural network thinks is cat-like too.(Or at least they pick it more often than if they were just selecting randomly in this seemingly nonsense task.) That is exactly “similar in some way”. Similar in which image they find more cat-like. That is the similarity.
I'm saying the implication is that there is a not-yet-understood deeper connection illustrated by this discovery. I don't think they tested for that and I don't think it's hinted at. There are lots of reasons why this trick would work on both humans and neural nets that would amount to effectively no similarities in structure or process otherwise. That isn't to say it's not related somehow, just that their experiment simply shows the outcome is the same, but doesn't indicate why it's the same.
Having said all of that, you're definitely correct to imply I haven't read the full study. And I'm happy to admit my original comment contains misunderstandings. Sorry about that.
My stats professor worked in cardiology and he shared a paper about n=1 or 2 studies: "If nobody dies, is everything OK?"
No wonder they didn't cite the actual numbers in this summary write-up.
> https://link.springer.com/article/10.3758/BF03206939
What is this link for? You linked an article published in 1993. The posted article is about a totally different one: https://www.nature.com/articles/s41467-023-40499-0
Edit: It seems you clicked the first link on their page. That lead you to the historical springer article and you are mistakenly describing that as if that is the current study. It is not, it is one they are citing as prior research.
For experiments 1 through 4, N was 38, 389, 396, and 389. The subjects were not undergrad psych students.
The article linked in the parent comment does not correspond to any experiment in the blog post or the Nature Comms paper.
> __Experiment 1__ included 38 participants with normal or corrected vision. Participants gave informed consent and were awarded reasonable compensation for their time and effort. Participants were recruited from our institute but were not involved in any projects with the research team. __Experiment 1 control__ (i.e., Experiment SI-5) included 50 participants recruited from an online rating platform. For __Experiments 2–5__, we performed psychophysics experiments using an online rating platform. In each experimental condition, approximately 100 participants were recruited to participate in the task (see Supplementary Table 18 for the exact number). No statistical method was used to predetermine the number of participants, but the sample size was decided to be comparable to that used in previous similar studies. Participants received compensation in the range of $8–$15 per hour based on the expected difficulty of the task. No sex or age information was gathered from the participants for all our studies. Our participants were all located in North America and were financially compensated for their participation. __We excluded participants if__ they were not engaged in the task, as assessed using randomly placed catch trials with an unambiguous answer (e.g., pairing an unperturbed dog image with a cat image and asking which image is more cat-like). If a participant failed one catch trial for Experiments 2, 3, and 5, or two catch trials for Experiment 4, the task automatically terminated and their data was not analyzed.
Parent's numbers are specifically drawn from Figure 3 caption. (Some text may not format correctly. Apologies if I didn't catch)
> a Participants are shown two perturbations of the same image, of true class T, and are asked to select the image which is more like an instance of some adversarial class A. The image pair remains visible until a choice is made. b One of the two choices is an adversarial perturbation that increases the probability of classifying the image as A, denoted A↑. Experiment 2: T = A; the second image is perturbed to be less A-like, denoted A↓. Experiment 3: T ≠ A; the second image is formed by adding a right-left flipped version of the adversarial perturbation, which controls for the magnitude of the perturbation while removing the image-to-perturbation correspondence. Experiment 4: T ≠ A; the second image is an adversarial perturbation toward a third class , denoted . c We show examples of adversarial images which empirically yielded human responses consistent with those of the ANN (indicated by the red box) for ϵ = 2 and 16, corresponding to the lowest and largest perturbation magnitudes used in these experiments. Example images in (a–c) are obtained from the Microsoft COCO dataset62 and OpenImages dataset63; images in (a, b, and c) left are used for illustration outside of our stimulus set due to license limitations. d Box plots (same convention as Fig. 2c) quantifying participant bias toward A↑ (where A = T for Experiment 2 and A ≠ T for Experiments 3 and 4), as a function of ϵ for four different conditions (each a different adversarial class A) collected from n=389 participants for Experiment 2 (cat n = 100, dog n = 100, bird n = 90, bottle n = 99), n = 396 participants for Experiment 3 (cat n = 96, dog n = 100, bird n = 101, bottle n = 99) and n = 389 independent participants for Experiment 4 (sheep vs chair n = 97, dog vs bottle n = 99, cat vs truck n = 98, elephant vs clock n = 94). The red points (with ± 1 SE bars) indicate the mean across conditions. The black dashed line indicates the performance of a random strategy that is insensitive to the adversarial perturbations.
File: /.../41467_2023_40499_Fig3_HTML.png.webp
RIFF HEADER:
File size: 671810
Chunk VP8 at offset 12, length 671798
Width: 2000
Height: 2255
Alpha: 0
Animation: 0
Format: Lossy (1)
No error detected.
So since they're lossy, maybe the subtly is lost?Edit: The image on the article itself is an SVG, containing 3 jpegs. So that's absolutely mangled in comparison to the paper's lossy images.
https://deepmind.google/api/blob/website/images/Figure0_svg....
Also these perturbation based adversarial attacks are often model specific. You take the model's gradient at each pixel and iteratively perturbate the image to make it more and more confident that it's e.g. a cat.
This is also why the “anti-facial recognition” shirts are silly. If they had adversarial noise at all, you likely won’t know about or have access to the model you’re trying to fool.
Remember, the ML model is objectively selecting cat with (very) high probability out of the entire corpus of possible responses. The human should be given the same range of possible responses to objectively establish bias. Since no human would say 'cat-like' for any of those images, it suggests a fairly large gap between human and machine perception. We've got a long road ahead of us.
Because the answer to that is simply “it looks like flovers in a vase”. There is no question about human’s ability to tell what the image is.
So much so that if you ask the humans to describe the images they would probably say something along the lines of “two identical images of the same flowers”.
So you would think if you ask them which one is more cat like they will shrug and pick one at random. Since it is a nonsense question. Yet people were able to pick up the manipulated image as more cat-like. Which means there is some signal they are able to pick up on.
> it suggests a fairly large gap between human and machine perception
Naturally. That is not at dispute, neither is it the subject of this study.
Thank you. Because that right there means that any bias being measured is one that's introduced by the researchers. Ergo the study is useless.
Yes. Intentionally.
> Ergo the study is useless.
No, it is not useless. It just not studying what you seem to think it does.
There was a previous very well estabilished finding that you can manipulate images such that they look basically identical to humans but they completely change classification for the neural network. This is an estabilished fact.
Have you heard about this? Are you aware of this? Because if you are not that would explain your confusion.
This research is not proving that phenomenon. It has been proven before. It follows up on it and investigates what that “basically identical” means.
There are two assumptions one can make:
1; The difference between the manipulated and the not manipulated image is so small humans won’t be able to tell which one is wich.
2; Since the manipulations appear to be neural network specific there is no reason to expect that a human will preceive them as the same class as the neural network. For all we know the humans might see the original image as more cat like, or the one we manipulated to look like a truck will appear more cat like to a human.
These are the two statements they wanted to investigate.
I agree we have a long road ahead of us, but you clearly do not understand the design of this experiment.
All this experiment measures is the impact of a priming effect.
> We excluded participants if they were not engaged in the task, as assessed using randomly placed catch trials with an unambiguous answer (e.g., pairing an unperturbed dog image with a cat image and asking which image is more cat-like). If a participant failed one catch trial for Experiments 2, 3, and 5, or two catch trials for Experiment 4, the task automatically terminated and their data was not analyzed.
But I fully agree, the experiments are poorly setup and they don't even have inter-correlating analysis. It's hard to tell if a force is going on, which there very well might be.
This study's result is implied by the stronger result: “a neural network's notion of a category sometimes resembles members of the category”. I'm sure a competent sketch artist could yield similar or better results, being able to take advantage of peculiarities of the human visual system. (In fact, that might be a good follow-up study: I might claim it if nobody else does.)
My subconscious pattern recognition for faces and such has always been weak, fwiw.
Maybe even the original authors would appreciate it.
Well, it turns out I can't mouse-draw, even when I'm tracing, but I've given it a go. It's supposed to be a cat walking from the right to the left of the image. The first layer has a better head, and the second layer has a better front leg and ears. Note that my lines obscure the recognisable features, so you'll have to switch the layer off to see them.
I traced dark shapes in the first one, and light shapes in the second. If you compare with the flowers layer, regions I identified as “better” in each trace correspond to amplification of the flowers image: the "good" leg outline in the second (light) layer lines up with the (light) stem of the flower, etc.
A million brainscans. A million images. Fed into an AI image gen. It's doable.
Many expected this, decades ago. But, also, many claimed that neural networks have nothing to do with the brain. I think we're slowly inching towards and understanding that we're the result of some fundamentals of information organization, and those fundamentals are realized in biology, rather than come from it. Those fundamentals are now showing themselves in silicon.
>when perturbed by a seemingly random pattern across the entire picture (middle), with the intensity magnified for illustrative purposes
I don't understand? The pattern is not "seemingly random", it is "seemingly chosen to have subtle cat-features". One sees the ears at the top of them image and face-like features below.
So, is it "we perturbed images to overlay cat-like features on a visual level that humans don't generally perceive but ML models were able to perceive; and then ML models perceived them"?
Can someone précis the results and why they're interesting because on the face of it this seems like a very obvious outcome?
Do I need to make a new year resolution to actually read the articles?
https://odysee.com/@NoVax:c/Hitchhikers-Guide-to-the-Galaxy-...
Unfortunately, the researchers will probably never follow up on their work and improve things. That's the sad story of most fun research, they just give up and forget about it once it's published and nobody else seems to want to do the work themselves. Though perhaps that's because most research doesn't actually have any potential for improvement or is false to begin with and the authors know it.
That is a very weird assumption.
> nobody else seems to want to do the work themselves
If you are so interested why don’t you try your hand at it? This particular research is super easy and accessible.
Ok, but the article doesn’t say what was the actual rate?
The effect strength on humans ranges from a few percent deviation of human judgements from chance for subtle adversarial perturbations (epsilon=2), to ~15% deviations of human judgement from chance for large magnitude perturbations in the largest magnitude experimental condition.
That is the label assigned to it in the dataset.
> Maybe it is just me on my phone but I might clarify it as a bouquet or flowers or something, but not a vase.
Ok. I assume you read the rest of the article. Does your observation change anything about the research findings?
I'm not an expert per se, but I think the issue at its core is that convolutional networks are trained to look at small features out of context, and tricking those smaller features detectors is possible without changing the overall structure of the image.
"They're the same--wait, is that a cat?!"
I mean, cats are small and cute, trucks are big and stinky.
And wouldn't the subject of the image also make the participants possibly lean to one of the answers? Flowers in a vase go well with a cat (both can be found in an apartment, for example). If the image would show the ISS, would more people tend to pick "truck", for example?
I understand that the concern could be something like "random perturbations of cat images make every image simultaneously less cat-like and hence more like anything else". My opinion is that Experiment 4 (making an image more cat-like or truck-like) covers concerns of this nature. Even if there are two random perturbations where one makes an image cat-like and the other more truck-like, it is completely arbitrary whether you label the perturbation as cat-like or truck-like (since they are randomly generated). That means, even if you measure a difference (and even if two random perturbations have larger differences than this construction!), you cannot control the direction. This method gives you a way to control it. Personally, I don't think this study is about measuring the influence over some baseline. It's about showing that you can indeed choose the direction of influence.
A broader way could be to add random noise then ask to pick the more X-like imagine and see how that correlates with the classifier probability for X.