Don't use text pixelation to redact sensitive information
bishopfox.com
bishopfox.com
> I’m actually not 100% sure why this happens (and sometimes doesn’t), but it’s an artifact of the rasterization process when text is rendered to screen.
This is a brilliant technique called subpixel rendering.[0] A "typical" computer screen's pixels are split into three columns of red, green and blue, instead of lighting up as a solid square. Using color fringing at non-pixel-aligned edges of characters can effectively triple the perceived horizontal resolution.
Wrong assumptions about the pixel layout will show up very badly, as shown in [1]. HiDPI displays (also mobile, not just 4K) get no real benefit from this and usually have a more complex subpixel layout anyway. I remember seeing old iPads usually showing awful fringes as subpixel rendering was enabled without considering the actual orientation of the display.
It also switched it off globally when the "magnifying glass" tool was active.
Most of these issues exist on other platforms as well. Try running a QT or GTK app in Mac, or an opposing widget kit app on Gnome or KDE.
Things have certainly gotten better (similar to HiDPI support, multi-monitor, etc) as developers have standardized or learned the tools, but an individual app can choose to do it’s own bespoke rendering on any platform.
I appreciate that reviewer every time I think about it.
Very fun project. Lots of problems out there.
PNG allows ASCII numbers, so flipping all digits to 0 creates a pixel which is graphically "masked" but leaks information about the original pixel: "000" means the value was larger than 99.
For a test in German class (my worst class), the teacher had just used tippex to remove some words and put them next to the text, and we had to fill them back in. I grabbed my ruler and measured all the sizes. There was 1 very long word, many medium sizes and a few smaller ones, but with this information and the context of the text for the first and last time I was able to get my first and last 10/10 in this class.
For some reason I just never trust the PDF tool (or human error on my end) actually redacting the info, even if I were to do a print to PDF.
I would absolutely not trust pdf not to leak metadata. Although now you risk metadata leak from the printer or scanner, which may or may not affect your threat model.
if you have the source document, redacting from the source (by actually removing and replacing with an appropriate placeholder, not obscuring, the content) and regenerate the static (e.g., PDF) version.
If you are working from print, I think scan and redact by digital replacement (not overlay or otherwise obscure) would be sufficient. Redact->print->scan probably helps somewhat (especially if the scan is low quality) if you are using a bad redaction method to start with, but why do that?
Of course, the artifacts introduced by printing and scanning (especially with contrast turned way up) gives it an air of legitimacy, although these can also be simulated.
Usually I'm in full control of the software myself so I just output X instead of the secret data.
This degrades quality and wastes paper and toner. There are software tools to convert PDF to raster graphics.
Recently discovered a manual forr some home appliance with a clear Word comment along with username, seems like slipped in when the manual was translated.
I've always wondered what some of these Melee players do as a real job and I guess in this case I found out by accident (in hindsight, the name of the company is in his Github bio, which I had never checked before).
Wanna redact text?
Put a solid bar over it. There, done. No takesies-backsies from that one.
And this is not so good advice if you're using software that supports layers and your solid bar is in another layer than what you're trying to hide.
Wanna redact text?
Delete it.
Redaction and Deletion are not equivalent. There are many instances where a visible redaction is required rather than a deletion.
...then really you should know how to flatten? Not understanding the tools you use for sensitive tasks like hiding/redacting data is a recipe for disaster.
> Delete it.
And this is not so good advice if you're using software that supports history and your deletion action is in another action that you can simply revert.
Replicating this process digitally is quite hard. Most pdf editors will put a square over the text, so that the original can be easily recovered.
2. Set Background Text Color to black
3. Put spaces in place of redacted text
4. Save
https://www.reddit.com/r/Wellthatsucks/comments/i1cdpl/the_d...
Redaction is usually a "Pro" feature in commercial PDF editors. People who use PDF's annotation feature to draw a big filled shape over the text/image to be redacted, don't know what they're doing and should be prepared for poor outcomes.
Adobe Acrobat and Foxit have "Pro" variants that redact properly, I'd love to know exactly how they do it. The open-source approach I'm familiar with, essentially converts the PDF to a flattened image and edits it to apply a colored shape over it. Example[1].
Although I just checked, MacOS Preview has a feature “redact” that (is supposed to) actually redact text. Well done!
That's not to say it's definitely not there, but it's at least something.
https://www.alfresco.com/blogs/corporate-news/redacting-pdfs...
>Redacting PDFs – What did the Manafort Lawyers get wrong? Date: May 18, 2020
>With Paul Manafort being released from Jail on May 13th, for those in the document space like Alfresco, it was worth revisiting the PDF redaction issue that surfaced during his trial. Back in 2019, lawyers representing Paul Manafort (a former lobbyist, political consultant, and lawyer, who chaired the Trump Presidential campaign team) filed a response to special counsel Robert Mueller’s claims that he violated his cooperation agreement by repeatedly lying to prosecutors. On pages five, six, seven and nine either the lawyers or the special council staffers attempted to redact sensitive passages.
>Although parts of the public version of this filing appeared to be redacted by black bars at first glance, it quickly became apparent that anyone with Adobe Acrobat, or other PDF viewing tool, or even browser-based viewing tools, could easily copy and paste the text that still existed under the redaction blocks to another document to simply reveal the passages that had been redacted. From the UK, a similar incident happened back in 2011 with the Ministry of Defence.
https://www.bbc.com/news/uk-13107413
>Internet mistake reveals UK nuclear submarine secrets
>The Ministry of Defence has admitted that secret information about the UK's nuclear powered submarines was made available on the internet by mistake.
>A technical error meant blacked-out parts of an online MoD report could be read by pasting into another document.
>Details were reported to include expert opinion on how well the fleet could cope with a catastrophic accident.
>The MoD said a secure version had now been published and it was working to stop such an incident happening again.
>Information also included measures used by the US Navy to protect its nuclear submarines, the Daily Star Sunday reported.
Well, mostly. But say you know the name that's redacted belongs to a small group (eg: US president since 1970 to today) - you could probably rule out (and mabye rule-in, determine) the redacted name, based on font-size, kerning etc in the document.
I wonder how fare a machine learning model could go, for longer reports - say 1000 pages with ~100 pages redacted - and the style of writing could be approximately inferred from the visible content - how many sentences/paragraphs could the AI fill in with some probability?
True, but sometimes there is no way around this, because redacting part of the text visually, so that it is clear where and how much was redacted, may be a requirement.
If there is no such requirement, one can always just replace the part of text like so:
This is the original text that I am going to redact now.
This is [REDACTED] now.Depending on what you're redacting you can eliminate this variable.
For example when I want to redact a piece of sensitive text in a video often times I'll make the solid black bar longer than the text being hidden. This way you can't infer the length of it based on the length of the bar. Of course this only works when you can extend the bar in such a way where it won't hide non-sensitive info that's important to see. In practice it works well, for example for redacting browser history just make the bar the entire width of the browser URL bar and for API keys or secrets often times the key exists as an env variable on its own line so extending the black bar is no problem.
I wouldn't count on that. A well-implemented print-to-PDF feature will try to avoid rasterising whenever possible.
If you want to turn your document (whether in PDF format or something else) into a raster image, there are proper tools for that.
It is a cool idea though, and text-shaped pixelation is much more satisfying than random noise. (And if someone's cheeky enough to decode it, they will be very disappointed ;)
Opaque gray boxes then?
[0] Ex. https://livebook.manning.com/book/programming-the-ti-83-plus...
frog: https://cdn.zappy.app/f97505ba92625a0e949aebcfa4220852.png
wikipedia: https://cdn.zappy.app/3b7d5cf750066633e40aeddda926f95f.png
It's not really targeted at the HN crowd (requires an account, etc), but the app is "Zappy" if folks are interested.
I have seen many a doc with history still included because they were using illustrator or something similar
I took down a ton of poorly redacted black boxes while modding /r/facepalm
In retrospect, a bot could do the job pretty well
The pixelized naughty bits censorship effect was more intended to cover up the humiliating fact that The Sims were not anatomically correct, for the benefit of The Sims own feelings and modesty, by implying that they were "fully functional" and had something to hide, not to prevent actual players from being shocked and offended and having heart attacks by being exposed to racy obscene visuals, because their actual junk that was censored was quite G-rated. (Or rather caste-rated.)
But when we later developed The Sims Online based on the original The Sims 1 code, its use of pseudo random numbers initially caused the parallel simulations that were running in lockstep on the client and headless server to diverge (causing terribly subtle hard-to-track-down bugs), because the headless server wasn't rendering the randomized pixelization effect but the client was, so we had to fix the client to use a separate user interface pseudo random number generator that didn't have any effect on the simulation's deterministic pseudo random number generator.
[4/6] The Sims 1 Beta clip ♦ "Dana takes a shower, Michael seeks relief" ♦ March 1999:
https://www.youtube.com/watch?v=ma5SYacJ7pQ
(You can see the shimmering while Michael holds still while taking a dump. This is an early pre-release so he doesn't actually take his pants off, so he's really just sitting down on the toilet and pooping his pants. Thank God that's censored! I think we may have actually shipped with that "bug", since there was no separate texture or mesh for the pants to swap out, and they could only be fully nude or fully clothed, so that bug was too hard to fix, closed as "works as designed", and they just had to crap in their pants.)
Will Wright on Sex at The Sims & Expansion Packs:
https://www.youtube.com/watch?v=DVtduPX5e-8
The other nasty bug involving pixelization that we did manage to fix before shipping, but that I unfortunately didn't save any video of, involved the maid NPC, who was originally programmed by a really brilliant summer intern, but had a few quirks:
A Sim would need to go potty, and walk into the bathroom, pixelate their body, and sit down on the toilet, then proceed to have a nice leisurely bowel movement in their trousers. In the process, the toilet would suddenly become dirty and clogged, which attracted the maid into the bathroom (this was before "privacy" was implemented).
She would then stroll over to toilet, whip out a plunger from "hammerspace" [1], and thrust it into the toilet between the pooping Sim's legs, and proceed to move it up and down vigorously by its wooden handle. The "Unnecessary Censorship" [2] strongly implied that the maid was performing a manual act of digital sex work. That little bug required quite a lot of SimAntics [3] programming to fix!
[1] Hammerspace: https://tvtropes.org/pmwiki/pmwiki.php/Main/Hammerspace
[2] Unnecessary Censorship: https://www.youtube.com/watch?v=6axflEqZbWU
[3] SimAntics: https://news.ycombinator.com/item?id=22987435 and https://simstek.fandom.com/wiki/SimAntics
Will Wright on Designing User Interfaces to Simulation Games (1996)
https://donhopkins.medium.com/designing-user-interfaces-to-s...
The Future of Content — Will Wright’s Spore Demo at GDC 3/11/2005
https://donhopkins.medium.com/the-future-of-content-will-wri...
The Soul of The Sims, by Will Wright
https://donhopkins.medium.com/the-soul-of-the-sims-by-will-w...
The Sims 1 Crowd Sitter
https://donhopkins.medium.com/the-sims-1-crowd-sitter-1f478b...
Dumbold Voting Machine for The Sims 1
https://donhopkins.medium.com/dumbold-voting-machine-for-the...
The Sims Pie Menus
https://donhopkins.medium.com/the-sims-pie-menus-49ca02a74da...
Automating The Sims Character Animation Pipeline with MaxScript
https://donhopkins.medium.com/automating-the-sims-character-...
Head Phaking, Phase 1
https://donhopkins.medium.com/from-don-hopkins-to-chris-trot...
The Sims Object Placement Tool
https://donhopkins.medium.com/the-sims-object-placement-tool...
And more documents at:
https://donhopkins.com/home/TheSims/
https://donhopkins.com/home/TheSimsDesignDocuments/
And some more videos:
The Sims Steering Committee - June 4 1998
https://www.youtube.com/watch?v=zC52jE60KjY
The Sims, Pie Menus, Edith Editing, and SimAntics Visual Programming Demo
https://www.youtube.com/watch?v=-exdu4ETscs
Demo of The Sims Transmogrifier, RugOMatic, ShowNTell, Simplifier and Slice City.
https://www.youtube.com/watch?v=Imu1v3GecB8
Transmogrify Self
https://www.youtube.com/watch?v=dsTbs7IL5EI
Speed Dating With Cupid
https://www.youtube.com/watch?v=YVUP9OXmHTM
Simprov Wedding Play Set
https://www.youtube.com/watch?v=Mwt5LJlrMe8
FreeTheSims Sims Character Animation ActiveX Control Demo
https://www.youtube.com/watch?v=Nzz0cFSmgiM
Also there's a great interview with Chris Trottier about "the toilet game", "tuned emergence", and "design by accretion", that I published on my old blog, which is still on archive.org:
https://web.archive.org/web/20160704065742/http://www.donhop...
>Sims Designer Chris Trottier on Tuned Emergence and Design by Accretion
>The Armchair Empire interviewed Chris Trottier, one of the designers of The Sims and The Sims Online. She touches on some important ideas, including "Tuned Emergence" and "Design by Accretion".
>Chris' honest analysis of how and why "the gameplay didn't come together until the months before the ship" is right on the mark, and that's the secret to the success of games like The Sims and SimCity.
>The essential element that was missing until the last minute was tuning: The approach to game design that Maxis brought to the table is called "Tuned Emergence" and "Design by Accretion". Before it was tuned, The Sims wasn't missing any structure or content, but it just wasn't balanced yet. But it's OK, because that's how it's supposed to work!
>In justifying their approach to The Sims, Maxis had to explain to EA that SimCity 2000 was not fun until 6 weeks before it shipped. But EA was not comfortable with that approach, which went against every rule in their play book. It required Will Wright's tremendous stamina to convince EA not to cancel The Sims, because according to EA's formula, it would never work.
>If a game isn't tuned, it's a drag, and you can't stand to play it for an hour. The Sims and SimCity were "designed by accretion": incrementally assembled together out of "a mass of separate components", like a planet forming out of a cloud of dust orbiting around star. They had to reach critical mass first, before they could even start down the road towards "Tuned Emergence", like life finally taking hold on the planet surface. Even then, they weren't fun until they were carefully tuned just before they shipped, like the renaissance of civilization suddenly developing science and technology. Before it was properly tuned, The Sims was called "the toilet game", for the obvious reason that there wasn't much else to do!
https://web.archive.org/web/20140720190605/http://www.armcha...
https://it.slashdot.org/story/07/01/07/1352242/blurring-imag...
which was republished in 2014 in Gizmodo:
https://gizmodo.com/why-you-should-never-use-pixelation-to-h...
See also related HN comments:
https://www.freecodecamp.org/news/lets-enhance-how-we-found-...
https://www.theguardian.com/world/2008/aug/15/thailand.inter...
You should also be careful with video, since it contains many frames it might be possible to reverse rougher pixelations or blurs than would be possible with a still image. There exist various superresolution algorithm that can extract high-resolution images from multiple low-resolution video frames.
Blurring is mathematically a convolutions product of an image that can be represented as a function f and a kernel g. The key fact is that the Fourier transform of f.g is the product of the Fourier transforms of f and g. So if you know g, or can guess what it is, you can solve for f knowing f.g and g, since everything is linear. (Some blurs might not be linear transforms but the usual gaussian blur is)
Information can be lost in two ways : 1/ quantization of the data due to storage in a limited number of bits, 2/ truncation of the blur at the image edges. So blurs cannot be fully reversed, but some information can be recovered. That might be enough to identify the information that was concealed in the first place. See the Wikipedia article on deconvolution for examples.
It’s just too easy to mess this stuff up.
Or just summarise the content yourself and hope there are no event/story/narrative based watermarks present.
Also, consider removing metadata from digital files using mat2:
A tool where you can choose the background color and the text color. The pixalation tool then overlays the blur effect with random characters.
I've used the original image data as a fill, but scaled down, so you get a mosaic effect. Then I randomise the tiles in that mosaic and then I blur. The result is that it seems like the original data was redacted but in reality the original data has been scrambled to such an extent that it can no longer be retrieved.
It's perfectly fine to use pixelation in the vast majority of scenarios where you are just trying to NOT draw attention to details that aren't part of your message. If someone gets nerd-sniped enough to actually sit down and spend a few hours deconvolving someone's name or email address, that's not the end of the world unless it's a very exceptional piece of information. Use your judgement.
In any case, manually blacking out rectangles to hide text looks ugly.
I mean if you don't care, why bother at all?
This way you keep the nicer aesthetics but you get the benefits of using the black box approach.
Probably overkill for many cases, but when I create any sort of content I like to keep it looking fresh.
The bottom line is that when you need to redact text, use black bars covering the whole text. Never use anything else.
That actually may not be enough if you're applying the black bar to compressed image data like JPEG because compression artifacts surrounding the black bar can be leaking information about the covert data.Though, web images now are being served in higher resolutions (200+ DPI) for "retina" displays, and scanned images are generally 300 DPI, in which case you'd be lucky even to get ascenders and descenders.
I'd be curious to give it a try though. If Facebook memes are any indication, many humans are totally oblivious to near-unreadable levels of artifacting.
JPEG breaks an image into 8x8 pixel blocks. Each of those blocks then has its information content reduced, so that it can be described in fewer bytes. (I.e., information is thrown away -- making JPEG "lossy", and producing visible artifacts.) This has the necessary side-effect that, when reconstituted, this 8×8 block now contains redundant information (if not, then the compression of that block was not lossy). This finally implies that at least some certain pixels of that block can be (at least partially) inferred from other pixels. That is, if lost, they can be recreated.
(It's helpful to understand also that JPEG does not encode each block on its own, but additionally factors out block commonalities into a central "dictionary".)
For the above to be useful to infer text hidden by a black box, requires:
(a) that the edges of the black box are not aligned to the 8×8 grid;
(b) that the relevant portions of text to be recovered lie near the edges of the black box (i.e., within the 8x8 blocks which straddle the edges); and
(c) that these blocks originally contained data of sufficient complexity, and/or deviating sufficiently from the rest of the image content, that the encoder decided to throw away sufficient information in these blocks to leave significant artifacts.
Finally, if the redacted image was re-encoded as a JPEG (or other lossy format), the re-encoding process must not have thrown away too much information in these blocks, else the redundant information will have been obscured and rendered all but useless for reconstituting the redacted information.
So, an easy way to avoid having redacted information extracted in this manner is simply, to ensure that your black boxes extend at least 8 pixels beyond the redacted text in each direction. (And also, to force the JPEG encoder not to re-use the dictionary from the original image, as information about the statistical distribution of block data could theoretically be extracted from that. Round-tripping through PNG is one way to force this additional safety measure.)
This still isn't 100% information-theoretic secure -- there's still residual information in artifacts elsewhere in the image about what patterns the original image's dictionary contained (which could be extracted with e.g. principal component analysis), which, when combined with a prior statistical distribution of the expected uncompressed content of the image, could leak some information about the portions which were redacted -- but I suspect the amount of information available via this channel to be vanishingly small.
Quick note: when giving examples of variable width and monospace fonts, the both look variable width. No monospace font is displayed (mobile Safari iOS 15.3.1)
I still do pixelize with huge pixel size when information is not that important. I think it better conveys that there was something there, and redacted areas look more organic.
I would not be so sure about cutting the area, don't forget to print it (to PDF) instead saving it.
1. Image is scaled down to create a mosaic effect
2. Pixels are moved around a random offset of x and y axis
3. Blurring is applied
Adding black bars over information will always be better but this does result in more smooth redaction that I think cannot be undone because of the randomisation step.
Only the original color data remains, but the detail is gone and each pixel position behind the redacted area is randomised/mixed. This randomisation step also overwrites pixel information so there is data loss as well. Because it's then blurred it looks like the info is blurred only. But if you'd be able to reverse the blur you'd end up with pixel noise. I find it hard to believe that that noise could be reverted.
Black bars may look very ugly in a video. Still, are video editing products recommending a process that has a high risk of leaking sensitive data? There might be reasons that attacking redaction in a video is harder than attacking redaction in a PDF. However, maybe it's actually easier in some cases, e.g., with several similar frames, the attack could take advantage of averaging across frames.
Unfortunately I don't see anything at https://hackerone.com/adobe that could get someone a bug bounty for researching this.
Print, cut with scissors, scan.
bishopfox: yangi !