Image Scaling Attacks
embracethered.com
embracethered.com
What the article doesn't mention, and the paper it links to probably mentions somewhere alongside so much irrelevant information that I couldn't find it yet, is whether this also works on some of the better scaling algorithms, and thus whether it's a "duh, OBVIOUSLY" or actually interesting research.
The blog post gives a cv2.resize example which seems to default to "bilinear", but I'm not sure what this means for downscaling, in particular for downscaling by a large factor.
I suspect that the key takeaway is "default downscaling methods are bad".
I have an image I crafted a long time ago which looks something like gray noise when you open it up, but when you downscale it, you see an image of Lt Cmdr Data from Star Trek. I wonder if I can dig it up.
The technique itself was not novel when I did it, a more sophisticated version involving embedded gamma values (which you can make quite large or small) was routinely used on image boards some ten or fifteen years ago.
http://www.ericbrasseur.org/gamma_dalai_lama.html
On my machine, both Firefox and Chrome display grey rectangles when scaling down. Why do the browsers get this wrong?
If its even possible at all. Sometimes users upload things like https://commons.wikimedia.org/wiki/File:“Declaration_of_vict...
You just have to go into a linear colorspace and use an area filter.
Bilinear interpolation is perfectly acceptable for zooming-in an image (making it larger by adding new pixel values). If you want to zoom-out, you have can still use bilinear interpolation, but of course you have to filter the image data beforehand to avoid aliasing.
The introduction of a mip chain + enabling mip mapping mitigates this, because when the scaling ratio is less than 50% the GPU's texture units will select lower mips to sample from, approximating a "correct" bilinear filter. This does also require generating mips with an appropriate algorithm - there are varying approaches to this, so I suspect it is possible to create attacks against mip chain generation as well.
Thankfully, quality-focused rendering libraries are generally not vulnerable to this, because users demand high-quality filtering. A high-quality bilinear filter will use various measures to ensure that it samples an appropriate number of points in order to provide a smooth result that matches expectations.
One other potential attack against applications relying on the GPU to filter textures is that if you can manually provide mip map data, you can use that to hide alternate texture data or otherwise manipulate the result of downscaling. As far as I know the only common formats that allow providing mip data are DDS and Basis, and DDS support in most software is nonexistent. Basis is an increasingly relevant format though and could potentially be a threat, but as a lossy format it poses unique challenges.
http://number-none.com/product/Mipmapping,%20Part%201/index....
http://number-none.com/product/Mipmapping,%20Part%202/index....
This is in essence a special version of sampling artifacts, aliasing artifacts. Anyone writing image processing software should already know about aliasing, the Nyquist theorem etc. Or, well, perhaps not in the current hype, where everyone is a computer vision expert who took one Keras tutorial...
Resizing with nearest neighbor or bilinear (ie ignoring aliasing) also hurts ML accuracy, so they better fix it even regardless of this specific "attack".
Also area interpolation still has some pretty terrible aliasing, since box kernels are terrible at filtering high frequencies.
And of course with downscaling you could still freely manipulate the downscaled image if you're allowed to use ridiculously high or low values, provided you knew the exact kernel used.
Area interp works very well in practice, it's more sophisticated than just a box filter on the input and sampling. It calculates the exact intersecting footprint sizes and computes a weighted average. Do you have examples where this causes aliasing and can show a better alternative?
Anything softer than area will help with those kind of issues (which is why the original https://en.wikipedia.org/wiki/Aliasing#/media/File:Moire_pat..., looks fine in most browsers even if your resize it). Bicubic tends to do better in this respect. It's a trade off though.
Now you could use pre-smoothing with a kernel and then resampling, but then we are talking about something else.
It's important to understand that interpolation happens in the source pixels, so it does not help when downscaling. Cubic tends to look nice, yes, but only when UPscaling.
> Any algorithm is vulnerable to image-scaling attacks if the ratior of pixels with highweight is small enough.
https://archive.org/stream/pocorgtfo15#page/n96/mode/1up
This isn't based on attacking scaling algorithms per se, but rather on the fact that most browsers honor the gAMA gamma setting in PNG file headers, while most image processing libraries don't and strip it when downscaling them.
The abuse potential for AI training exists here too, but both attacks are a bit of a stretch.
... but yeah, it's a screech as the application in the real world, seems to be a really specific case to work.
Internet Explorer had this feature that if you CTRL+A page contents, it would overlay images with 1px grid to indicate selection. If you got your pattern right, the hidden image would appear. This is essentially the same effect, but on steroids.
Even if this method is not feasible as an attack vector, at the very least it looks like a very practical way to share information that otherwise would be censored or restricted–all the more so if the hidden image data can be encrypted, which may make it impossible to detect.
On the other hand, I know nothing about steganography and I'm talking out of my arse, so maybe current steganography methods are much more powerful.
> No, that's what's so cool about it. I explain it in more detail in the video, but basically because the videos are created using 1-bit color images, it makes it easy to retrieve data without having to worry about how YouTube changes the video.
There's a video explanation here: https://www.youtube.com/watch?v=yu_ZIr0q5rU&feature=youtu.be
Source code here: https://github.com/AlfredoSequeida/fvid/
An example here: https://www.youtube.com/watch?v=NzZDFxM5Coo
I can easily imagine how someone could use this for nefarious purposes.
first there's lossy compression, which means that there's no guarantee your injected pixels survive the encoding pass.
Then there's the additional hurdle of motion vectors, which will most likely be misaligned between the original video and the injected one.
This would result in hard to predict artefacts after encoding.
Finally, each decoder handles scaling slightly differently, so even if your embedded video trick works on one software/hardware decoder, it might fail on another (sometimes even depending on just the version or additional settings/filters being enabled).
https://twitter.com/GalacticFurball/status/13197659867911577...
Probably hard to get it working in every environment, but if you know what you are up against, it might be possible ;-)
It's still cool though