Exploiting aCropalypse: Recovering truncated PNGs
da.vidbuchanan.co.uk
da.vidbuchanan.co.uk
I hope that people who host forums, image boards, chat applications, etc., will delete or fix potentially vulnerable images before anyone uses them maliciously.
One way to repair a vulnerable image is to use `optipng -fix`.
Reading about this silent API change makes me feel like I'm losing braincells. What's going on with the processes behind Android's development?
Google's engineers working on Android at a system level regularly break basic functionality in the "userspace"*. Google's engineers working on Android apps get early access to the Android versions and work through the resulting bugs, bubbling them back up until they get fixed.
(*userspace being used loosely here, it's all userspace vs being in a kernel, but it's interfaces that are implemented at the OS level and consumed at the app level)
Like Google is large enough that I'm sure someone will take offense to implying that such questionable engineering takes place there, but this isn't a story I've heard just once. People working on apps that are part of every GMS enabled Android image have confirmed this completely bizarre setup on multiple separate occasions
You're relying on your internal teams' relatively outlay to catch contract breaking. And sure in this case internal is Google, meaning there are some pretty widely used apps acting as filters... but relative to the millions of apps out there, they're not likely to catch all the regressions.
They must have automated testing, but at the point where it's just accepted that things will break and your own engineers regularly have to "convince" your OS team that they broke things, you know something is wrong.
This is clearly unacceptable, but I've seen so much worse.
Even worse when they don't do this and just flat out admit they broke a usecase intentionally and label it Won't Fix because the team that implemented the breaking change is also the team that triages the issue. See logcat access being completely broken in Android 13: https://www.reddit.com/r/tasker/comments/wpt39b/dev_warning_...
This is absolutely horrifying.
Yeah, especially in this case, due to changing defaults and similar-but-differently-behaving APIs.
Defaults really suck sometimes. But so does not having any. And so many things can become security issues when used just so.
:/
That's why shipping a new API requires a lot of time investment in the design of the API: once an API is shipped you can't just change the behavior dramatically.
It does leak info from inside the box due to jpg compression artifacts.
Here is the proof of concept I quickly wrote to show why this isn't safe : https://github.com/unrealwill/jpguncrop
Typically that apply to things like redacted pdf lossily compressed as image, for which you have already a few candidates words (and you can bruteforce). You try them one at a time and see whether the compression artifacts match.
The uncropping algorithm is pretty straightforward in theory : remove jpg artifacts, and fill the cropped region with candidate image portion x, compress and compare the produced artifacts to the cropped image artifacts, try a neighbor candidate x+eps*N(0,1) and optimize (aka random search).
The artifacts are related to Fourier coefficients so the distance between artifacts isn't too irregular.
The remove jpg artifacts, can range from really simple to really hard depending on the class of problem you have.
If the background image is something digitally generated (like here in our example of red and blue circles), or a pdf file you can get the uncompressed version without mistakes.
If the background image is something like a photo then you need a little finesse : you need to estimate the compression level, then run a neural network outside the cropped regions (4 different images : above, below, left and right of the cropped region) that remove the artifacts (something like https://vanceai.com/jpeg-artifact-removal/ but finetune to your specific compression level), so you can estimate the artifacts, then you search for the image inside the region (eventually with a neural net prior but that increase the probability of data hallucination) such that the jpg artifacts of the compressed reassembled image, is close to the jpg artifacts of your cropped image.
Do you have a proof of concept of the technique you're describing? Otherwise I remain skeptical.
One important piece overlooked by not doing an actual demonstration of blind reconstruction is lossy compression is naturally not a 1:1 mapping of original sources to compressed sources. If it were it'd be a lossless compression! This means your theorized recovery process ends early as it assumes that the first match is the only match and therefore the original. Had the claims of been fully demonstrated instead of assumed to be the same as recreation this would have been accounted for.
In practice JPEG blocks represent an 8x8x3=192 byte/1536 bit original source. From that we take the lossy compression and get some small number of representing bits (which is encoder and settings dependent) back, let's say 154 (1/10th size) for discussion. Of that some number of bits is going to be used to accurately encode the black square, let's say 1/2 just to be generous to the extraneous noise and let's also be generous and say the re-encode from original JPEG to obfuscated JPEG was nearly lossless (i.e. "100" quality) to help the numbers further. We're now at 154/2 = 77 bits representing 1536, in the generous case. 2^1536/2^77 = 1.5 * 10^439 possible original matches. Even assuming the image is largely losslessly compressible the numbers still don't come out in good favor. For most cases, like the pictures example, this rules out finding the original being practical without some additional guidance that hasn't been demonstrated either (and maybe there is some other guidance! It's just not shown what that would be or why we should assume it exist).
What would be a really good demo is focusing on something minimal (e.g. text instead of images) and regenerating the source material without knowing it before hand. E.g. it's safe to assume a typeface of a document that was screenshot and later obfuscated and now all you need to do is show that one understandable string of text has artifacts that match a short whiteout/blackout on the document. This is still an extremely hard problem but information theory tells us it might be (not that it is) at least reasonably possible.
I am just showing that jpg compression cast a digital shadow on itself that can be outside of the cropped area, which many people find unintuitive.
Matching the shape of this digital shadow (which is not very small (all the highlighted pixel region contain some info) ), combined with an educated prior on the missing portion (that must provide a way to represent the hidden portion in less bits than you can get from the shape of the shadow) , can help you recover the data.
If it's so easy and already done why not show it actually reconstructing real examples like that page. The reality is there is no existing well known approach for recovering data in the exact way you laid out. Similar approaches sure but they rely on different data sources (such as downsampled versions of text) instead of artifacts. That it is feasible in the easy case does not mean it is feasible in the hard case, why should the two be expected to hold the exact same amount of information about the source? Even if they do the method of reconstruction will surely be different.
> I am just showing that jpg compression cast a digital shadow on itself that can be outside of the cropped area, which many people find unintuitive.
That I agree 100%, and it's a great thing to show, it's just a bit overreaching to then name that "jpguncrop" with a description referring to aCropalypse instead. Maybe it could lead to something like that but what's in the repository does not match the name or description.
> Matching the shape of this digital shadow (which is not very small (all the highlighted pixel region contain some info) ), combined with an educated prior on the missing portion (that must provide a way to represent the hidden portion in less bits than you can get from the shape of the shadow) , can help you recover the data.
Great, you have a theory - demonstrate! I'll be the first to comment on how clever, persevering, and well executed the work was if you post a thread showing it works. Until that point though it isn't so just because you think it would work.
Here is a great test case https://i.imgur.com/sAEpfQ6.jpg. The font size should be small enough that data leaks out of the obfuscated area from the characters. I'm optimistic this blind reconstruction case is feasible given the additional constraints compared to the circle test but it would probably take a decent chunk of new work.
It'll be fiddly to exploit (like the "undredacter" you linked), but I think you could get a text-recovery PoC going. Try to craft an "ideal scenario" (for an attacker - i.e. ideal censor-bar placement and font size), see if you can exploit it in practice, and build from there.
"The end result is that the image file is opened without the O_TRUNC flag, so that when the cropped image is written, the original image is not truncated. If the new image file is smaller, the end of the original is left behind."
RIP
Some implementations of EXIF stripping might help, but it's not guarenteed.
Edit: I made my own. I can confirm that the exif chunk was not stripped. https://cdn.discordapp.com/attachments/541730746805649476/10...
But yes, a more sensibly coded EXIF stripper would deserialise and reserialise. Unfortunately I am no longer able to assume that programmers will behave sensibly.
Edit: Also, the PNGs generated by Markup don't contain EXIF in the first place, so an EXIF stripper could reasonably decide that no changes are necessary at all.
As you noted above Discord doesn’t sanitize PNGs. This exposes a failing on their end as well, as large services taking input from users should sanitize images to protect both senders and recipients.
This isn't a metadata issue.
An underlying IO library changed its behavior so that instead of truncating a file when opened with the "w" mode (as fopen and similar have always done, and this API did originally), it left the old data there. If the edited image is smaller than the original file, then the tail of the original image is left in the file. There is enough information to just directly decompress that data and so recover the pixel data from the end of the image.
You're not necessarily recovering the edited image data, just whatever happens to be at the end of the image. If you are unlucky (or lucky depending on PoV) the trailing data still contains the pixel data from the original image - in principle the risk is proportional to how close to the bottom right of the image the edits were (depending on image format).
The bug is you say "write to this file" which is meant to erase the existing file if such exists, but the underlying library either had a serious regression, or intentionally broke API compatibility, and changed the behavior to not erase the existing data. Your exif stripping + reserialization would write the new data down and the trailing data from the original file would still be present: e.g. exactly what is happening in this bug.
No amount of processing in memory, no amount of reserialization, no amount of data filtering prevents this bug. The bug occurs at the point of IO, because the IO is meant to have erased the original file, and it did not, so if you write fewer bytes to the destination file than were present in file being overwritten the tail of the overwritten file remains and is leaked.
To make it very clear that this is not an error in processing the image: if you opened "image1.png" (or whatever format), edit it, and then saved it over a different file that already exists, say "image2.png", and then send image2.png to someone, this bug will allow the recipient to extract the trailing data for the original image2.png, it would not show any information about the original image1.png.
Not that I'd want to maintain custody of such a dataset...
…
>This also applies to the "Snip & Sketch" tool in Windows 10.
Edit: I find myself wryly weighing this against the ongoing unleashing of LLMs upon the world. Both have shades of clever people prioritizing being and demonstrating clever at the cost of... other stuff. On the bright side, it is distracting me from facepalming at the underlying Pixel bug.
All the tool seems to do is just read out whatever comes after the end of the PNG and then supply the missing data to construct an image that can be rendered.
The only thing I can think of that would have made a real difference is to send a tool to fix the images to all image hosting platforms in advance. But which ones do you trust?
Some people would just lose interest if there isn’t an easy tool immediately available, and also it would give potential victims or image hosts more time to fix or delete vulnerable pics.
What if we just... thought twice about making it this easy?
Everything after that is fair game.
Image edit metadata also seems like an incredibly useful feature. Do we just strip it as well?
Do you have a specific concern to warrant your comment?
Is this possibly a helpful feature or is it really just a terrible hack/bug that has no practical use holding on to a sort of edit history inside a PNG?
I would love a way to track some level of history in a commonly supported image format (but of course being aware of needing to strip it when appropriate)
So this behaviour is likely unfit to be used as a feature, but could in some cases be used as a clever hack to, in a way, preserve some edited (cropped) data.
You sound like you're probably fairly smart but I suspect you're rushing this one and commenting before you've properly grasped the topic.