A tool for recovering passwords from pixelized screenshots
github.com
github.com
To me, it seems like a lot of wasted time and effort and potential arms race between encoders and decoders and the risk of being exposed when you could just put a big black box over whatever you wish to obscure.
The only benefit I can think of is legitimacy. Having something blurred there suggests that there was actually something there.
First, if you add black bars on your own and don't use professional redaction features of software, you might miss the OCR text layer of the PDF, or the bar might be added as a separate object entirely which means it can also be removed later on.
Second, if you don't use monospace text, the width of the text you are redacting will reveal information about it. That's why monospace fonts are so commonly used in the intelligence community for example.
Third, if you just add a black bar to a screenshot, there might be residual values of the text left in adjacent seemingly white portions of the image, but they might not be entirely white due to compression effects. Better you run it through a filter before publishing.
Looking at you, Preview on macOS.
If you do use monospace text, width reveals the exact redacted character count. If you don't use monospace text, it constrains the possible contents in a more complicated way, but it leaks information either way.
Such a great trick. Just write random letters or even random sentences over your handwriting perhaps 2 or 3 times, then heavily cross it out and it's impossible to recover.
I do a lot of screencast videos and what I typically do is put a black rectangle well past the point where my secret text really ends.
If my secret is:
mysupersecretpassword
You might end up seeing a black bar of: ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
What I'm most afraid of is accidentally having literally 1 frame of video where the bar isn't present. This is super easy to do if you're just scrolling through a timeline while editing. You need to really dive in and move frame by frame to ensure you get everything. I usually extend it a few frames in either direction just to be safe too.However, glitches in editing software can still rarely occur, so for very confidential footage I first censor it, render it out in an intermediary format, check the render to make sure it's 100% correctly censored and then use that rendered footage to continute editing.
Also, mixing/messing with frame rates is an easy way to shoot yourself in the foot.
I'm currently looking at BlackMagic's DaVinci Resolve, which has a quite well featured video and audio editing package for free. But 'all' the free version does is edit video clips. You have to pay for additional functionality. Premiere Elements, in contrast, has lots of useful things like creating/handling non-video assets such as text for titles and credits, loads of (sometimes cheesy) clip transitions, etc etc. Resolve may be okay if you hate Adobe.
You'll need something more sophisticated if you are targeting non-standard playback devices, or want to participate in collaborative workflows, or grade raw sensor clips (assuming you can get this data off your video device). As a real power user you might also want to look at ffmpeg, something I simply haven't had time to get my head round.
Again, YMMV.
[Edit] I consider myself an 'advanced' stills photographer, who is much happier using manual rather than auto settings. When I started with video, however, I quickly discovered how much I didn't know. For example, what is the relationship between the shutter speed of the camera and the frame rate of the video? This was helpfully ignored in the camera's manual, and took some googling to find out. Video autofocus is also pants, on the D850, for the type of wildlife and action photography I like (it always refocuses just as the bird lands). So I'm teaching myself to use manual focusing, just like they do in the movies. A video tripod head was also essential so I could pan, elevate and focus with only two hands rather than three.
The free version lets you create titles, animations, transitions and you get access to their after effects comparable tool for fancy animations as well as color grading and audio / video editing.
It's a very reasonable free offering. I use it all the time to edit podcasts. I want to use it for editing video too but I have issues with it adding artifacts to the rendered videos, I think it's because it doesn't like my 6 year old GTX 750ti.
Took a while to convince people to not use a bit of gaussian blur, because it's insecure. Well, ready for round 2...
One that will shuffle around or maybe just randomly derive the pixelated blocks if they're within a certain color threshold of similarity?
I still generally like pixelation over drawing a box for aesthetic purposes.
From a security perspective it has zero advantages, but I guess some people like the aesthetics?
>To me, it seems like a lot of wasted time and effort and potential arms race between encoders and decoders and the risk of being exposed when you could just put a big black box over whatever you wish to obscure.
There doesn't need to be an arms race, one can just erase whatever needed to be hid, put fake random text over it and then pixelate that. Ie., a prettier "black box" but still the same core method as a black box. In principle doesn't even need to be any more work, it'd be easy enough to throw together a simple "pixelate erase" script that'd take a selection, apply average solid color and random text over it, then pixelate the result all in one automatated step.
The problem is if some people think any modification of the original information is ever good enough. Anything based on the original could leak information, so need to erase and then apply any desired aesthetics afterward.
I think this whole pixelation thing is just for making it look nice. I, personally, would just use something like asterisks without any pixelation.
Or replace the original password for "password" and leave it at that.
Averaging the color leaks information. To avoid leaking info, destroy the pixels as the first step. Get colors from outside the selection.
Applying it to an HN comment the result would be a solid average of #F6F6EF (beige background) and #290027 (text), so a solid darker beige with some random black text over, all pixelized. How can any of this be used to recover the original text?
Say for example, the password field had one of those password strength colour indicators that subtly changed the background or font colour to indicate password strength, and the images were captured and obfuscated at that point. Information useful to an attacker would now be embedded in the colour average.
The attacker then before trying to hash a candidate password they can first calculate the average color of it to check if it even remotely matches, which can be much faster than the hash function.
Average color could be a rough predictor of password length too, depending on circumstances.
This makes sense for purely aesthetic reasons because the pixelized block will not stand out against the combination of background + text around it. A white page with black text would have a gray fuzzy area where it's pixelized.
If a weighted average is used and you can determine the fill factor of the text inside the box that would be some information leakage, as small as it may be.
https://twitter.com/Phthalaldehyde/status/133471474074092748...
This is either a typo or I want to know more about it.
https://www.google.com/amp/s/amp.knowyourmeme.com/memes/reco...
Which is in and of itself a security risk. Not only can the length of the redaction give an estimate on the length of the password, but for variable width fonts, could even rule out many passwords using pixel measurements between the text on either side.
- Someone makes the background color black in Word and saves as PDF. Text is still selectable - Someone draws a black rectangle over the text in Word, text is still searchable / rectangle can be removed with any simple PDF editor
[1] https://9to5mac.com/2018/03/13/ios-markup-reveal-redact-sens...
For a similar reason, designers use "lorem ipsum" placeholder text rather than all-white or all-black placeholders when mocking up a layout
It looks less jarring when pixelated compared to being blackened out.
I think one of the reasons people use blurring/pixelation is that it looks nicer than black boxes.
Wondering if it would make sense to build a tool that renders a pixelated version of lorem ipsum or something like that. It would be secure while also looking good.
These kinds of mistakes show up constantly. Another one that often works is when people draw black bars over words they are often no 100% opaque and you can chuck them in to gimp, drag the levels around until the original word is exposed.
Cannot find source.
> Two French hackers used their computer skills to reconstruct a blurred-out code on TV and claim bitcoins worth $1,000 (£760).
I think the hackers could have made more money if they got real jobs doing valuable work.
Reference: https://boingboing.net/2007/10/08/untwirling-photo-of.html
It seems to me that this would only work in very rare situations. It needs some kind of pattern that is created with the same font. This pattern gets really large if you add a lot of chars (i.e. a-z already makes it huge).
I tried this with a really simple example (numbers only, cutting exact frame of pixelated image, same font for pattern) and it didn't produce any useful result.
TBH that doesn't look all that impressive. It's not a "throw pixelated image in here and get the unpixelated result" tool.
(It's still of course valid to make the point that pixelation can be insecure. It's of course very much possible that much better such tools - likely by throwing in some ML - are possible and may even exist in the hands of people who won't share them on github.)
If the goal is to de-pixelize some text in a public website (login form, username, phone, etc.) this is as simple as right click > inspect > input the sequence > screenshot.
I think US courts and lawyers are finally starting to learn this.
This way the redacted PDF won't have any metadata from the original (e.g. bookmarks) nor risk any weired PDF redaction bugs.
The downside is that your redacted PDF is (probably) larger than the original and isn't accessible. You can mitigate this by using OCR but it still isn't perfect and has a high chance of messing up tables or any special text placement.
Some fancy combination of encodings (e.g. using JBIG2 except for images) can probably help with the filesize.
---------
The best way to handle sensitive information is probably to use codenames (for people, places, etc) so you can share the original documents without worrying so much.
This is really the collision between designers and programmers. Due to aesthetics, designers try and cover text up by using a blur. Programmers see this from a functional stand point, they see some security issues. Programmers act to prove their methodology is superior and that is the story of how this repo came into existence.
- You know the exact parameters used to render the text
- You can render new text with the exact same parameters
- The pixelated image hasn’t been ruined by color quantization or other destructive compression
It would also be interesting to see how well it worked if there were differences in font rendering or compression as you say - I wonder if you it might still be close enough to make a partial match in some cases.
https://www.linkedin.com/pulse/recovering-passwords-from-pix...
To preserve the same aesthetics (a blurred username or identifier that doesn't call attention to itself) but would make it impossible to match against?
E.g. in Photoshop, the "Filter > Pixelate > Mosaic" option should have a checkbox called "Secure" or "Security noise". Or there should be a separate filter called "Pixelate > Redact". Ideally it would use some intelligence to figure out the size of characters/symbols in the selected area and automatically figure out the right combination of pixel size and noise.
A screen from Super Mario Bros is pixelated. That's what the art is supposed to look like.
A screen from Broken Age running in Retro Mode is pixelized. The art isn't actually intended to look like that, and it's just for nostalgic effect.
Still cool edge case.
$ 7z a 12345.7z passwords.txt
$ cat 12345.7z >> cats.jpg
Try: https://i.imgur.com/McqO2eX.jpgThe author used this property to create the de-pixelizing algorithm. A neural network could do it, but it would be overkill.