Mark Felt-Tipped: Uncovering top-secret information by counting pixels
matthi.coffee
matthi.coffee
Good catch. Interesting, but obvious once you think about it. Not really a danger for properly chose passwords/passphrases (the search space would still be too large for brute forcing) but it would highlight where a brute force attempt is worth trying.
I might have to add an extra (hidden) input to all our authentication pages programatically filled with random characters up to a length of 1024 minute the length of the entered password.
IMO if you don't need the entropy, it's probably easier to pad with some kind of "clearly not legal" character which is still visible when debugging, such as newlines.
You almost certainly don't, but my default for any security related value is "properly random" except where the randomness itself might give clues. That way around is less likely to result in a "kicking oneself" situation in the future!
In the case of hashed passwords, suppose the user entered "foo" but you stored hash("foo42156"). Now the user is locked out, and there's no way for you to fix it on their next login attempt, because you have no way of knowing how much of it was "real" anymore.
In contrast, a deterministic system (like "pad with newlines to 256 bytes") allows you to take their next login attempt, validate it under the older method, and "upgrade" the hash to the correct de-junked version.
It's not just passwords either: The issue of corruption also applies to all variable-length non-hashed sensitive data you might apply this scheme to. For example, security-questions, e-mail addresses, financial account numbers, etc.
In this case though the extra padding would be in a separate field that only exists to control the length of the POST request body - nothing should be looking at it in a way that would allow it in to corrupt other data.
password = "foo2556019562042" # Example limit of 16 chars
password_real_len = 3
Article: https://www.washingtonpost.com/news/politics/wp/2018/02/25/w...
Handy GIF: https://img.washingtonpost.com/pbox.php?url=https://www.wash...
"By September 2016, the FBI had opened investigations into four members of Trump’s campaign team. The Democratic memo says the information compiled by Steele into his infamous “dossier” of 17 raw intelligence reports didn’t get to the FBI’s counterintelligence team until the middle of September. By that point, we can conclude thanks to a sloppy redaction (noted by former intelligence officer Matt Tait) and an unredacted footnote that Page, Papadopoulos, former Trump campaign chairman Paul Manafort and Michael Flynn, who would go on to be Trump’s national security adviser, were all already under investigation."
Matt Tait links to: https://twitter.com/pwnallthethings/status/96752319618133606...
“Simpson said Steele first shared his concerns with the FBI during the first week of July 2016 and in a subsequent meeting with the Rome official two months later when Steele provided the official ‘a full briefing’ of his findings”
https://www.usatoday.com/story/news/politics/2018/01/09/doss...
https://www.politico.com/story/2018/02/24/democratic-memo-go...
This statement could be perfectly true, while it's also perfectly true Steele met with the FBI in July, and had multiple other channels to provide information to various other FBI departments. It doesn't particularly matter exactly how the investigation started. If you've read the dossier and now knowing what we know about how it came to be, it's pretty disgusting.
"FBI officials indicated that Steele himself was not advised that the work he was doing was on behalf of the Clinton campaign."
Now this is something I hadn't heard before! That is absolutely shocking that we're supposed to believe this ex-Spy was in the dark about who was paying him?
So the reasonable question that follows is: was the sloppiness intentional, or a leak?
[1] https://firstlook.org/theintercept/article/2014/05/19/data-p...
Also worth mentioning that just like the Secret Service has an ink database on all the printer types in the world, the NSA is supposed to have a database of what different keyboards sound like. This means that simply by recording the sound of you typing, they can infer keystrokes / characters. Obviously the easiest way to record this is by hacking your phone, which is right next to you.
- Assign a prior probability of letter frequencies based on a corpus of the language text you're analyzing.
- Separate the different keystrokes in the audio file into a series of times between keystrokes. (i.e, have your program recognize one keystroke as distinct from another)
- Based on the subtle timing differences between keystrokes, assign various lengths to different timings.
- Using the prior probability of the letter frequencies, assign the different time lengths to different characters based on their frequencies.
- You now have a straightforward mapping between the distance between two keystrokes and the character typed, which should allow you to decode the typed text.
Further calibration can probably be had by considering a word dictionary and using fuzzy matching to detect how often words are decoded incorrectly and what the correct decoding would be.
Of course, if you know or suspect that you are under such surveillance, you could try and alter your typing cadence, e.g. by switching to hunt-and-peck.
Maybe a random layout generator that would mask this effect and only cost typing speed for the extra defense?
In this example, the POS would be an adjective, and since the subject noun is plural, it would be more likely the adjective is a number
That makes short redactions more dangerous for declassification than entire paragraphs as a general rule because you have no context to start from, but that's probably common sense to people doing redactions.
Some kind of Reed–Solomon type encoding in the typography that would allow retrieving the whole document even after it has been redacted and copied.
e.g. Go from "XXXX people connected with.." to "[REDACTED] people connected with". So long as it's noted that redaction happened and one doesn't have the source document to glean things from, that would probably be sufficient to block the things the author of this article is doing.
What I'm wondering, though, is if this is permissible. Being able to see how much of a document was redacted gives us some important context:
1. We know we're looking at the "original" report. This might not be the case with an approved transcription of some sort. 2. It's easy for a layman to get a sense of how much information is being withheld.
Knowing what and how much was redacted seems like it's pretty important for maintaining trust and transparency with an organization that has legitimate reasons to withhold some amount of information.
That's a good idea, I might do a quick Show HN about that.
"iF YOu wAnT tO SEe yOUR deMOCRacY aGAIn lEAvE 5 miLLioN dOLlArs uNDer THE bRidGe AT ..."
"Hey boss, we got a subscriber to Mad Bomber Monthly, Torture World, and BYTE."
(flipFlipFLip) "One of these three guys?"
We will throw in another 5 if you also take your closet-case VP with you!
redaction by wingding could be a feature
Due to new regulations all documents must be padded .It's a classic packing problem - searching for phrases of exact pixel width, where each of the unknown number of letters contributes a different number of pixels.
I think it is also unlikely given the space that follows too which would read more naturally if it was name + title, but it could be if the page was just off too much (I think the article here does a better job aligning than this guy did from looking at the beginning of the sentence.)
I wonder how brute-forceable it is. Given a blank with length of half a line, can I just throw 20-30 combination of the alphabet and punctuation, and figure out which combinations match, and the anagram for it?
Actually, brute-forcing using words from a dictionary (English words plus names of the involved people (Americans, and obviously Russians)) would make it go even faster.
If not, it still seems tractable using dictionary words.
Yes.
I would be interested in seeing how many "matches" digits would yield
Why don’t they do this?
How big is the secret they're hiding? Is it of unlimited or limited length? Your technique would obscure the sense of scale.
However, out of concern for individual workers in the bureaucracy, I would support a tag on such stories to warn them before clicking.
EDIT: On reading the other comments here, I don't think even a tag is warranted for this story as no leaking has occurred.
Insufficient care has been taken in redacting parts of the document. Responsibility for that falls to whoever redacted the document, not for whoever was able to deduce the redacted content.
Such people do not deserve special workarounds by the rest of society because of their terribly poor judgement.
Since its never authorized to access classified information from an unclassified network, it makes sense to block access to classified materials from unclassified networks.
I was thinking that too, but it characterizes Carter Page as a "former campaign foreign policy advisor". The other names may have similar designations, making that a hopeless game.
We do know, from an unredacted footnote, that one of them was Michael Flynn. So that makes 2/4.
You could iterate the strategy further by creating a full font from the existing text, identifying spacing and etc for each next letter/character, calculate for word wrap lengths, and get it even closer.
https://gist.github.com/anonymous/37200394d975e1a86ee9d19bef...