Unredacter: Never use pixelation as a redaction technique
github.com
github.com
As a sibling comment to yours also points out, there are plenty of examples to the contrary - models taught to recognise text from indistinct images. We’ve all been contributing to this for years via Captchas, and now we’ve got to the stage where our iPhones apply the concept locally to text found in our photo libraries.
I ask because what was there would have affected the original compression algorithm. So there might be some signal left.
Let’s say you’re covering a red pixel with a black pixel. There’s no way that red pixel will in any way “survive JPEG compression” because the compression happens on the raw pixel data; the raw pixel data no longer has a red pixel.
They're saying that before encoding the red dot could have been X pixels wide, and to the naked eye it would seem that after encoding it's still only X pixels wide... but due to compression artifacts it may now actually be affecting an area X+Y pixels wide. So only covering X pixels could leak Y pixels of information.
It honestly sounds a bit paranoid to me, but it's definitely possible.
This is trivially untrue for JPEG; JPEG often stores chroma info as 4:2:0, that is at a quarter the resolution of luminance info. In this common case, covering a single red pixel in the decoded form only removes a quarter of the chroma information from its presence (and this is easy to not notice if, for example, the other three pixels that share its chrominance are nearly black).
So, if you want to redact a rectangle between pixels [ 12, 4 ] – [ 34, 18 ], redact [ 8, 0 ] – [ 40, 24 ] instead, this way the black rectangle will cover complete blocks.
The only issue I see is meta tags and embedded thumbnails, which some cameras do. Stripping EXIF is always required.
JPEG compression causes ringing artifacts that can extend out past the pixels you're trying to redact, up to the next block boundary. This is a problem if what you're trying to redact is already a JPEG, rather than a pristine uncompressed image.
This.
I made a response to the parent question that goes into the practical side of this in a bit more detail, but this is a much better one-sentence summary of /how/ the information leaks.
I have also seen people try to redact by drawing a black block on a new layer in Photoshop, for example, and then exporting as PDF. They're unaware that the layers are being exported in the PDF and the block can be removed. Always flatten your document completely before export.
Whether that's enough to decode the redacted information will depend on how many pixels were redacted and other intricacies of JPEG.
This is not too dissimilar to compression oracle attacks, such as CRIME (https://en.wikipedia.org/wiki/CRIME), an infamous vulnerability in HTTPS.
https://en.wikipedia.org/wiki/Rescue_of_Giuliana_Sgrena
https://en.wikipedia.org/wiki/Rescue_of_Giuliana_Sgrena#Rele...
Consider the following case: You have limited previously knowledge of the unredacted document; for example, you know that it was a pure black-and-white document that was compressed as a standard color JPEG (not too unusual). Suppose further that in the block extending from (16, 16) to (24, 24) (exclusive), you wish to redact the single pixel at 23, 23, presumedly as the upper left corner of a larger black bar. Finally, suppose that the post-redacted image is losslessly compressed; this simplifies the argument, but is not strictly necessary.
In the way we've phrased this problem, there are only two values for the original color of the redacted pixel; it was black, or white. The rest of the pixels in the block are known in their encoded form (unredacted), but also known in their pre-encoded form (pure black or pure white, while JPEG compression will put them somewhere into dark gray or light gray, respectively). With all known pixels set to their pre-encoded form, each possible value of the redacted pixel can be considered, using (known, deduced, or guessed) JPEG compression parameters to regenerate the compressed block. Whichever value of the redacted pixel leads to the observed encoded values of the unredacted pixels is the correct value of the redacted pixel.
(In practice, the problem might be less constrained; and you might have multiple but less than a block of pixels redacted; these make it harder to make use of the leaked information, but do not change that some information has indeed leaked.)
Programs that appear to black out a section of image, but actually leak information.
(Of course, you don't actually have to know how to pronounce something in order to write it! But if you decide what to write by saying it to yourself, you might find yourself instinctively reaching for some alternative turn of phrase.)
which provokes ye olde zen koan: if a joke is made in a pixelated font whose typeface will never be known, is it still funny? :)
idk, is it? :^) https://imgur.com/1OvGeAF
>One of the first things you might notice is that it has a curious bit of coloring in it. What gives?
>I’m actually not 100% sure why this happens
It's subpixel rendering which exploits the way pixels are laid out physically on a monitor (1 logical pixel is actually three physical subpixels: red, green, blue). It produces crisper texts.
If you're nearsighted taking your glasses off can blur the pixelation of a face so that you can get a pretty good idea of the person's appearance, especially if you are just trying to guess if it's really a specific person you are already familiar with.
Now, pair that with a Prosopagnosia disorder... Right back where you started?
There are AI algorithms that 'enhance' blurred faces, however, the output 'looks' like real human face, but in all likelihood, different from the ground truth (i.e. the real face that got blurred). Therefore caution must be exercised in advocating such algorithms (e.g. in court of law)
Whereas this algorithm is 'just' enumerating the text, generating blur and finding the closest match to the original blur.
I doubt the US TLAs have it, because the software they have is often rubbish generated by a cartel of low skill providers, and frequently 10 years behind the times. But I bet the Chinese Ministry of State Security has it. So so many dissidents to keep track of, to identify after they have moved overseas and apply pressure to.
Don't use text pixelation to redact sensitive information - https://news.ycombinator.com/item?id=30350626 - Feb 2022 (163 comments)
What you mention would be a random permutation, which seems vialirt ot vercore given its small search space.
Photo identification is done with a limited search space of “tiles” that the image is decomposed into, for convolutional NNs.
If it's secure enough for intelligence agencies, it's secure enough for me.
In more realistic applications you’d have to deal with things like text that is not exactly aligned with pixel boundaries, and different anti-aliasing methods.
E.g. a page to do fold recognition on, or even to figure out what software (e.g. Gmail) and to look up what font the software uses.
I've wanted to build something that would if nothing else run through a list of guesses (assuming the font is the same as surrounding text) and see if any of them could match size-wise, but not sure of an easy way to deal with the PDF part of it.
If you aren't doing this so often that you need automation, you could sidestep that issue by just taking screenshots at 400% zoom or so and accurately measure how many pixels each letter (a-z) takes, as well as how many pixels a space is (the censored part might be into the spacing on each side of the word), and measure how many pixels the gap is between the words surrounding the censored part, then
for word in wordlist do
# Start by accounting for the spaces
wordsize = charwidths[' '] * 2
for char in word do
wordsize += charwidths[char]
done
if wordsize == gap_size then
print("Possible word: " + word)
endif
done
Probably want to do the gap_size +/- 1 or so, but that's how I'd approach this for a given document. A starting point for a wordlist on many linux systems could be `/usr/share/dict/words`.If you need redaction, black it out completely.
before the above the author made a version which was free (but old link is now dead):
https://github.com/Y-Vladimir/SmartDeblur/downloads
LEO's have access to a Commercial tool called Amped Five that is said to be very good at it:
Also the example seems to go through one letter at a time, once the pixels become larger it might be required to cycle through entire words.
That's a clever idea for speeding up this process actually. Instead of guessing a-z for every position, only guess valid word completions. If it doesn't end up giving a good score or the human judges it to be nonsensical, it can fall back to a-z guessing for that word.
A bit similar to my hangman solver (https://lucgommans.nl/p/hangman-solver/), which looks for the only words still possible with the given letters already known, but simpler because you only need prefix matching.
I know you meant it for very strong blurs, and there it is not actually an advantage because you need to go through thousands of words before guessing one (instead of 26×5≃130 guesses, assuming an average word length of ~5), but yeah there you'd have no other choice.
But now just in response to the title... does anybody ever use pixelation as a redaction technique?
I feel like I've only ever seen pixelation for censoring nudity or a brand name or face or something. Actual redaction of text truly meant to be kept secret, I've only ever seen as black bars.
Are there notable cases where people have genuinely tried to redact something secret using pixelation?
I've also seen people try to redact sensitive information by poorly scribbling it out using the pencil tool on their phone, leaving enough parts of letters visible to guess what was originally there, or try to "black out" text using a brush that wasn't fully opaque, allowing it to be revealed by adjusting the contrast. Basically, a lot of people don't know how to safely redact stuff.
I don't think pixelation is commonly used because most people probably don't know how to use it, but my Samsung phone's built-in image editor also has a pixelation feature, so I honestly expect to see it pop up more often in the future.
A colored square is easy — and destructively removes what was there, by setting that part of the image to a chosen color.
But just imagine the face of someone getting Rickrolled in this fashion.