Lessons learned from cracking 2 million LinkedIn passwords
community.qualys.com
community.qualys.com
cat /usr/share/dict/words|egrep -v "é|'s$|[Åå]|[Øø]"|shuf --random-source=/dev/random -n4
This uses the dictionary /usr/share/dict/words and skips all the words containing characters like é, å, ø and all those ending in 's. The resulting word list has 72,940 words in it. Then it chooses 4 random words from this dictionary and prints them to the screen. This gives a password with about 65 bits of entropy.By adding another word, thus creating a 5-word passphrase, a botnet capable of checking 1,000 trillion passwords per second would spend, on average, 1600 years cracking away before it would find the correct passphrase.
Here are some example 4-word passphrases produced using this method:
poetically archaisms accept constrictors
leukemia shuttlecocked checkout benevolently
climactic gyrate dynamical predominates
massage beef Concords recliners
These are surprisingly easy to remember. I use a 7-word passphrase for the most important things and it didn't take me more than a day or two to learn it. massage beef Concords recliners
Damn you.Edit: also, for anyone wanting something like the above technique, I recommend Diceware.
I haven't seen any passphrase advocates say that.
The ones I've seen all say "use 4-5 random dictionary words". Come up with a meaning for the phrase, and it'll be easy to remember.
egrep -v "é|'s$|[Åå]|[Øø]" /usr/share/dict/words|shuf --random-source=/dev/random -n4http://en.wikipedia.org/wiki/Cat_%28Unix%29#Useless_use_of_c...
LC_ALL=C egrep '^[[:lower:]]{4,8}$' /usr/share/dict/words |
shuf --random-source=/dev/random -n4 |
fmt
That's picking from ~35,000 words, which I think is still good enough but avoids ending up with "interpreters incredibility disciplinary constitutionality". (I need the C locale to stop this old GNU grep being painfully slow.) 72940^4 = 2.8304992 × 10^19
To put it in perspective, a 14-character password using only lower case English alphabet letters as individual tokens already beats this: 26^14 = 6.45099747 × 10^19The key takeaway here is that every word adds another 16.15 bits (assuming good random source and no non-random decisions by the user), whereas another character adds only 4.7 bits. I'd argue that the effort to remember another 4 random characters (to reach those 16 bits) is far more than the one to remember another random word. We're quite good with words, you know :)
WeampE6quaph (Weamp-E-SIX-quaph)
2hoov2Klypfo (TWO-hoov-TWO-Klyp-fo)
GicutOj8 (Gic-ut-Oj-EIGHT)
HegEmWydwev5 (Heg-Em-Wyd-wev-FIVE)
Tegdijetyik4 (Teg-dij-et-yik-FOUR)
Fon7ochry (Fon-SEVEN-och-ry)
I usually take the password and type it about 100 times to see how it feels, subtly changing any characters that feel awkward to type. This gives me a password that is very fast and natural to type (less likely to have errors), but still has a lot of randomness and obeys all the stupid password rules.Even relatively high-stakes companies like banks and credit card companies make obscenely stupid mistakes when it comes to security. For example, there was a case fairly recently where you could log into your Citibank account and change the account number in the GET query string, and you'd instantly have access to anybody else's account. Given that they're capable of that type of idiocy, all it takes is for one mental giant to decide that encrypting your password is better than hashing it, and you're vulnerable.
Malicious behavior isn't the only thing to watch out for. By doing business with the outside world, we're putting ourselves at the mercy of complete morons every day. If you use a different passphrase for every account, then you can at least limit your risk to one service.
function mkpw () {
if (( $# == 0 )) then
head /dev/urandom | uuencode -m - | sed -n 2p | cut -c1-${1:-12}
else
head /dev/urandom | uuencode -m - | sed -n 2p | cut -c1-${1:-$1}
fi
}
By default it generates an alphanumeric string of length 12. Given an integer argument n it generates an alphanumeric string of length n. tr -dc '!-~' </dev/urandom | head -c ${1-12}; echo
Change to A-Za-z0-9, etc., to suit.So this is what I've been wondering about the current "best practice" to use long passphrases. How are those really any stronger than any other "rule" based password, the "rule" being that they are likely constructed of words and phrases from human language.
Would the passphrase "My first car was a 1972 Monte Carlo" really be harder to crack (once the cracking tools are adapted) than a random 8 character password?
And technically, if you didn't know the format of the password, and you were just trying to get a random 11 character password, that would take a long time to crack. There are (roughly) 94 character that you could safely use for your password pretty much universally on any website...
94^11 = 5.06x10^21 which means if your computer can generate 2 million hashes a second it would take: 80 million years to crack a truly random 11 character password.
Passphrases are stupidly insecure unless you throw enough randomness in it.
ex(quotes included): "My Phone Number is `(123)546-8794!!!`"
The memorization of that password would work much better than a simple passphrase like that.
I.E. the actual password would be:
"That's a battery staple. Correct!"
And I don't believe that people will easily be able to crack that even with the minimal randomness that has been put in with current techniques. Sure if natural language cracking becomes popular you may have to become a little more creative like using a made up word or name or a number but even your example if no one knows what your password is: "My Phone number is (123) 546-8794."
should be sufficient for a very hard to crack password. And again is many times better than a simple dictionary passphrase with a few words combined.Wait, what?
2,048^4 == 2^44 == 17,592,186,044,416
At 2 million hashes/second it would still take 101 [edit: actually, on average, 50] days to find this password, if it was unsalted. Perhaps if you had spent a few years of supercomputer time to generate some massive rainbow tables, you might be able to discover it quickly, but absent the need for your linkedIn password to be resistant to attacks from a nation state, you'd be pretty safe with such a password for a while.
It's entirely unclear how you came to the conclusion that it could be discovered in "under a minute" with a passphrase dictionary attack.
Say my algorithm is to pick the password "1" * 1000 (that's the character 1 repeated 1000 times) and also pretend that 90% of the sites didn't have stupid limits and it was a valid password. It's certainly a long password. The time it would take to brute force it by testing all possible strings in order of increasing length is an unimaginable number. It's not on the scale of the universe - not on the scale of a million universes either.
But now let's say that this "the more characters the better" became a universal truth and everyone jumped on the same bandwagon and did the same quick hack of having 1000 1s. Suddenly, we're all screwed, because the algorithm "pick 1000 ones" is staggeringly weak. In fact, it provides no protection at all - the attacker already knows your password.
The true measure of security measures is not how long they last when no one knows about them - it's how long they last when everybody knows. "Pick 10 random symbols" will last for a while. "Pick 'password'", not even a second.
Where does "pick a meaningful English sentence" fall on the grand scale? That's one incredibly hard question to answer. It's also bloody difficult to break, for reasons of generating sentences, not password entropy.
Did you read the article? It describes exactly what a possible attacker does. And it's not "start with short passwords".
There's only two options:
- Use a really random password string, from a non-broken random generator
- Do something nobody else does
The latter only works if you can stop yourself from bragging about it on public fora. Which is why one of the best pieces of advice for secure passphrases is to include something really, really embarrassing, horrible, shameful, completely unfit for print and absolutely boring. Especially don't use a funny quip or play on words, don't try to be clever, there ought to be no audience to appreciate it. And if at all possible it shouldn't even look like a password.
(kinda OT) I read that advice many years ago, and I don't understand why Julian Assange did not take it to heart. Remember when that Guardian journalist wrote his book and published the passphrase to that AES encrypted data dump (because the nitwit assumed the AES passphrase would be automatically invalidated after a few hours ...), it was something like "a diplomatic history from <date>" with some random uppercasing, special characters, etc. It would have been pretty strong, except it was WAY too clever and typical-super-secret-password-looking to use for the sort of hypersensitive data Assange was carrying about. If he had simply picked some terribly bad and misspelled slashfic involving Martin Luther King, a dead baby and pres. Nixon--like Spider Jerusalem would've done--no way the Guardian journalist would have published that, anywhere.
I come from down-south India and I talk Malayalam.
For a typical password, each character can be one of around 92 characters, depending on what rules are in place - 26 lowercase letters, 26 uppercase letter, 10 digits, and ~32 special characters on the keyboard (I may have miscounted). Other characters could be used, but these are going to be the most common.
This means that your 8 character password can have about 100^8 possibilities. To put that into more familiar, and more easily comparable terms, that's 1x10^16 password possibilities.
According to Oxford Dictionaries, "The Second Edition of the 20-volume Oxford English Dictionary contains full entries for 171,476 words in current use." This means that, without reducing that space, a four word passphrase would have about 8.6x10^20 possibilities.
Admittedly, there are some massive problems here. The most obvious of which is the fact that most of those 171k words aren't words a normal person would use. For this to be a valid analysis, you would have to believe that the average person would pick a passphrase like "gastroenteritis jurisprudence algorithm aberration", which is clearly ridiculous. Also, most people would, like your example, use a grammatically correct sentence. The possible combinations would be pretty severely reduced in that case.
Now, more combinations are introduced by capitalization, punctuation, and the introduction of "numeric words", like the year 1972 in your example, but I have no idea how to account for that.
In either case, the average person is going to have a much easier time remembering "My first car was a 1972 Monte Carlo" than they will remembering "8gj2;hg^".
Oh how I wish my bank and mortgage lender would let me choose easy-to-remember passwords like that.
There are many more short words than long words, thus a person would be very unlucky to pull out that passphrase.
But what if you reduce the space? Instead of using a dictionary with about 175,000 words, why not use the Diceware list, which has only 7776 words? None of them are over 6 letters long (I think.) A few words are numbers; or have special characters.
Because many websites won't allow you to use a diceware passphrase you'd use a good password safe with a long diceware passphrase. You'd then let the safe generate random passwords for you.
But change to something like "my first grandma was a 1927 haircut" and you're likely to future-proof it significantly.
Let's look at a password like "My first car was a 1972 Monte Carlo". The password is 35 chars, 3 upper case, 6 special (spaces), and 4 numbers. The key space is all upper and lowercase english letters, all numbers, and all special characters. That's a key space of 95 characters, over 35 places. Objectively, there are 1.66 x 10^69 possible combinations. Given that the LinkedIn password crackers are slowed down at about 9 chars it seems like you're incredibly secure. But let's assume the attacker knows something about your password structure. Let's say they know that you use words (many people do, so it's a reasonable guess). Let's also assume that for numbers the attacker knows that years are popular for password numbers. Now instead of 35 chars, your password has 7 words and a date. We've changed the key space from 95 to about 100,000. (The exact number of words there are is a tricky number to pin down, but crackers have some good data on what the most popular ones are.) As for the date, there are really only a couple hundred interesting numbers, including all dates from this and last century, as well as common patterns.
Password strength is (key depth) ^ (key length). An uninformed attacker has 1.66 x 10^69 possible combinations (95^35), while an informed attacker has roughly 1.0 x 10^40 possible combinations (100,000^8). Obviously, the less an attacker knows (or can guess) about your password structure, the better chances your password has against being cracked.
Now, you asked about your password versus a random 8 char password. Let's take a "strong" password like "1~qQ%57h" This password also has upper and lowercase letters, numbers, and symbols. We can assume that there is nothing predictable about this password for this exercise. The password strength is 95^8, or 6.6 x 10^15, obviously much lower than the longer sentence, even if the attacker knows the sentence is 7 words and a date.
Now remember, our passwords are being matched against human crackers attempting to guess the ways our passwords are most likely put together. For now, most passwords are 6-12 characters. In fact, most websites only allow passwords of these kinds, so it makes the most sense for crackers to go after these passwords. But it's still an arms race. If we assume that webmasters see the light and allow (or enforce) long, sentence-like passwords, the crackers will adjust. It's plausible I think that 5-10 years from now, we'll see articles like this one that use sentence structure syntax as an attack method.
Until we discover and implement a better system that obsoletes passwords, the best we can really do is have long, complex, and unique passwords for everywhere we go, and have a system to manage them for us. I believe that something like LastPass or KeePass are the way to go for now.
*Disclaimer: This was written on a groggy Sunday morning. Do not rely on my calculations. Do not use any of the examples as passwords. Do please check my work.
I guess if someone stole their database it would be impossible to know your real password, but still...
Or am I missing something here?
Yes it's to save space ...
No, wait.
It's so they don't use all the CPU power ...
No, not that either.
It's because the programmer didn't want to use their braincells.
Yeah, that would be it.
Rainbow tables have a degree of freedom: the function that maps hashes back to passwords. You should try and pick a password that that function will never generate. To get that, do something unique. Good options, I think, are including a foreign language word (neither English nor your native language, nor the site's language), reversing a word or a syllable inside it, and made up words that have Hamming distance greater than two to any other 'obvious' word.
Short (<= 8 characters) passwords, I think, are bad choices for that reason, even if they consist of ASCII gibberish.
Disclaimer: I have never looked what kind of code commonly used rainbow tables use.
Besides, rainbow tables are supposed to be pointless because everyone's supposed to be using salt with their passwords...
Making an intelligent phrase will affect the distribution of initials, but even something commonplace like "the quick brown zebra jumped over the mooon" or tqbzjotm hits the less frequent letters like q and z. It won't be completely random, but it's going to cover way more of the 8 letter space than words are.
Pretty much every site with a login facility has an "I forgot my password" option where you put in your username or email address, and it sends you a link to reset your password. This is effectively a second form of authentication - the ability to receive email at that address implies the ability to log in to that account.
So what about an authentication mechanism that works as follows:
1. You type in your username
2. The site emails you a one-time authentication token (as part of a link)
3. You click on that link and then you're logged into the site
Of course, there are a few obvious problems with this: it's a bit cumbersome to have to do for every login, email is unencrypted, and message reception uses a pull-based mechanism.
So I could envisage a standard, incorporated into browsers, as follows:
1. When you first launch your browser, you log into an authentication server S, supplying your password (either manually, or automatically via a saved password)
2. When you want to log in to a site, you type in your username, and the site sends an authentication token to the server S
3. S sends a push notification to your browser with the authentication token
4. Your browser passes this token to the site, and you're logged in
This way, your (hashed + salted) password need only be stored on server S (in the first example, this corresponds to your email server). This means that apart from S, none of the sites you use need to store any password information at all.
I'm sure this basic idea has been implemented previously in other contexts. Why are we not using for all our web logins?
Surely this needs to be promoted more heavily to website developers, especially those behind major sites like linkedin. I'd never heard of it myself and I'm sure a lot of others haven't either.
If some of the bigger sites started adopting it, awareness would grow very quickly and we could greatly reduce the risks involved with every website storing password info for all their users.
And as if that wasn't bad enough (which it is) it is a single point of failure both in terms of security and availability.
You could also have multiple accounts with different S servers, e.g. one for work and another personal use.
I agree with the single point of failure regarding availability - if your authentication server is down, you won't be able to log into anything. Though we already have the single point of failure with the existing system, in that once someone has your email password, they can obtain password reset messages from any site that you've registered on with that account.
You could also have multiple accounts with different S servers, e.g. one for work and another personal use.
If I was forced to use such a service I'd make a service that made it easy to automatically create one "S-server account" for each "real" account and continue to use passwords for those accounts as if nothing had happened.
In practice, BrowserID doesn't solve anything for me - at the cost of reduced security, availability and integrity - as well as forcing me put trust in a third party.
There is a huge difference between my mail server and the S server. If someone uses my mail to reset passwords I will notice, since my credentials won't work anymore. Also there are different levels of security, I value my mail account more than say my account on hacker news. Which I haven't even entrusted with my mail-address - love that you don't have to supply even a fake one and considering that I don't forget my password (or allow anyone to hijack my session) I can't possibly gain anything from supplying it.
Which is the key point, rather than me not trusting ycombinator there is just no incentive for me to supply it - so why should I? Maybe ycombinator gets hacked and my mail gets leaked, I might thus end up with spam - no need to take that miniscule risk when there is nothing to gain. Just as I see no reason to link independent accounts together with a service such as BrowserID.
The whole point of a password is that you can remember it. The moment you need software to store and retrieve passwords, you're better off using asymmetric cryptography. That said, I really dislike the idea that it's the only way to achieve security. I would really like to see more discussions and propositions for solving this problem.
Also, do you really think that storing private keys on RFID would be a wise choice? I would put that in the "stupid as fuck"-category, just above storing your keys on a USB-stick, as they are both easy to copy and duplicate, and with RFID I could do it remotely.
A real solution to all of these problems would be for people to stop reusing passwords. You don't really need the account passwords when you basically have access to all the data there anyway.
Slow hash function certainly help, but I think we also need something that goes beyond straightforward cryptography to address authentication issues. Something that redefines the rules of the game to be more human-friendly and less computer-friendly.
But then again, I never even heard the question being phrase this way: what do we want from "program-less" authentication and what we can use to achieve it.
With a password manager why bother restricting any password to anything less than the maximum? Gmail's password limit is 100 characters so I did that and any other account they are maxed out. Add to that extra authentication and also change them at least once every six month at minimum.
The problem is my most valuable account, my bank, is stuck 15 years in the past.
Or, just make it where websites HAVE to state somewhere how they are storing the credentials. It's shocking how many places still use plain text, or encryption and store the key in the database..
It's pathetic that a major company like LinkedIn is simply storing credentials with a SHA1 hash. At LEAST use a really good salt...
example: http://csrc.nist.gov/publications/drafts/800-118/draft-sp800...
I like this idea. Like a Surgeon General's Warning for the web. I wouldn't want the government making specific laws about hashing, but requiring transparency and disclosure about how data is stored would be useful in a variety of ways.
OTOH, it may be worth it. It's shocking that LinkedIn could be so negligent, especially after high-profile screwups like gawker.
Only the script kiddies. The ones you have to worry about have bots and automated scans that can figure that stuff out in an instant.
Yeah, unbelievably shocking that such an advanced web company as LinkedIn could be so negligent. Amateur bitcoin sites, social media sites, venerable Web 1.0 ones like Last.fm don't surprise me much, but LinkedIn? WTF.
There are some newer password management/hashing tools. I've stuck with this one both because it works for me and I know and trust the authors, a group at Stanford.
print math.log(62) / math.log(2) * 6
35.72 bits
That's easy to crack. Also, keep in mind that humans don't select chars randomly. So the bit-strength of these passwords was probably closer to 20 bits. I cracked 2.5 million with an old cpu and JtR within a few hours.
www.keepass.info
I must say, just finding the password reset function on some of these forums and less popular sites is a beast. Also, I was shocked by the number of 10 and 12 char limitations I hit.
The current version of Chrome would allow someone to, within a few clicks, grab a pile of passwords.
Here's another scenario: Your mother takes her laptop to be repaired/updated. She uses Chrome. The entire repair shop has easy, unencumbered access to all of her passwords and logins.
Similar scenario: Computer goes to IT guy where you work for repairs/updates. He now has any and all of your passwords and logins with no effort.
My point is that for all this talk about security it seems really dumb for a prominent player (any prominent player) to not take extra steps to ensure that our valuable data is secure within reason. With LinkedIn the problem is, at the very least, the lack of anything beyond SHA-1 to protect passwords. Bad idea. In the browser case, it seems to me that, unless the intent is to provide a browser used only by those like us who understand and are very aware of security issues, it might just be a good idea to put in a few things that will make it harder for curious eyes or the 16 year old at the repair shop to grab all of your login data.
I don't propose nor do I expect perfection or absolute security, but what Chrome does today is, in my opinion, at the very least irresponsible. The uninformed user has NO IDEA WHATSOEVER that a huge security hole exists in their browser. Maybe we need to stop thinking in our terms and focus on mom, dad, uncle or grandma. When you first install Chrome you should, at the very least, see a screen telling you about security and the options you might have. I think that a master passwords would most-definitely serve a purpose in the case of "innocent" peeking. Yes, with pro's all bets are off. It's only a matter of time until someone tracks identity theft to the lack of browser security and they sue the fuck out of the browser publisher.
With a USB stick and one click anyone can install malware that would give complete control of the computer to the user remotely.
> Computer goes to IT guy where you work for repairs/updates.
IT repair guys generally need admin access to the computer and will have all the time in the world to install any number of malware for remote access.
> but what Chrome does today is, in my opinion, at the very least irresponsible
For Chrome to add a master password would be irresponsible because it would give users the illusion of security they don't have. All OSes already have password protection against innocent peeking with user accounts and the ability to lock your computer when you walk away.
where x is the length of the password you want.