Poul-Henning Kamp: LinkedIn Password Leak? Salt Their Hide
queue.acm.org
queue.acm.org
This is a change that would have to be added to the current project backlog, specified and designed, developed, and implemented. Selling this means making a compelling case that salting and changing our hashes would actually solve a problem for us and our clients. My sense is that this is the case, but articulating the case in a unassailable way is still something that needs work.
The most compelling case would be for our clients to demand this as part of their security requirements for our systems. This sword cuts two ways, and a number of our existing password policies are clearly based on well-intended but somewhat misguided client-based requirements. The sane thing is to get good requirements.
Absent that, the question becomes: what is the threat, what is the risk model, what is the mitigation, and what benefits does that mitigation buy us and our clients.
The risk as I see it is disclosure of our user authentication hashes (thank Krell we're not storing cleartext passwords ... at least not there).
Leaking unsalted hashes means that both rainbow tables can be applied against the known hashes, and that duplicate hash instances (hence: duplicate passwords) can be determined and targeted for rainbow/brute force attacks.
Leaking non-bcrypt hashes means that brute-forcing is cheap. At some estimates, 3.3 billion keys per second on $1000 of hardware for MD5, roughly half that for SHA1 (http://www.extremetech.com/computing/84314-how-to-secure-you...).
A successful attack would gain user access, and might gain access to user information (of varying but largely low sensitivity) and be able to impersonate the user for communications purposes. Some collections of user data might be valuable for contact/communications/social-engineering purposes.
The biggest risk would be for users sharing keys among several services. As a fair number of our clients are corporate, and it's fairly well known that corporate password policies are often even more grossly weak than individuals', the likelihood of compromised passwords being used to access other user accounts in some instances is fairly high.
The question remains: how large are any of these risks?
What does salting and bcrypting buy in way of protection?
My read is that, on the technical side:
- salting hides common passwords within our userbase, and renders rainbow tables useless. Weak passwords are somewhat better protected.
- bcrypt makes the costs of brute-forcing passwords markedly more expensive. Very, very weak passwords could still be cracked, but we're talking on the order of searching through perhaps a few millions of keys -- 4-character alphanumeric mixed-case passwords would be at risk.
- checking proposed (or entered) passwords against a known set of common passwords -- even just a few tens of thousands of the most common ones -- would further reduce low-hanging fruit. Ideally I'd like to see a publicly available corpus of all known passwords, to be used to exclude duplicates.
But again, the question becomes, what demonstrable benefits does this present to us and our clients? How do I make the case?
Information leaks are common: a backup tape gets FedExed to the wrong address, file sharing gets accidentally turned on, a Russian hacker finds a security hole in your machine while scanning millions of machines, some idiot puts the password database on a laptop and loses it. These sorts of problems are constantly making the headlines.
If you have bcrypt-style password encryption, such leaks are a nuisance and embarrassment.
If you do not have password encryption, the leak recipient can easily impersonate any and all users. They can control your system, create false communication, cause industrial equipment to destroy itself, send harassing messages, conduct financial fraud, and so forth.
The cost to use password encryption is a little engineering labor, the return on investment is a substantial reduction in risk.
One is the perceived fear of looking incompetent in front of your users/clients. For which I feel the appropriate response is "we'll look a lot more competent if we mitigate the risks of such an event than if we don't, regardless of whether or not it happens".
But really, the big one is simply: can you justify the engineering/product cost of this change on the basis of a material business benefit to us and our clients?
Our backups management is pretty solid, with backups encrypted, and even DB systems using on-disk at-rest encryption via an ecryptfs tool.
You did raise the valid point of sensitivity of identity data among some of our clients. While the general case is that PII (personally identifying information) disclosures would largely be embarrassing but not harmful, there are cases in which harm, or even life-threatening risks could arise.
I'm leaning to your conclusion but I'm looking to be able to quantify that more robustly.
And as I noted in my original question: if we were getting pressure from our clients on this, the case would be far easier to make. Market rules.
These guys did a nice example: http://howsecureismypassword.net/
Lists are available from various sources. Here's a good page http://www.skullsecurity.org/wiki/index.php/Passwords
Then, it would need to become popular enough for users to start to recognise it and look out for it when signing up. Even if most users don't have any idea what it's about, plenty of the more technically inclined users would, and they tend to be the early adopters anyway...
The idea is to add a bit of pressure to services to store passwords correctly (similar to how users look for the green SSL bar when doing important stuff online), and providing some transparency to the users who care about this.
Honestly, I see it as almost self-evident that user would never ever learn this.
But more importantly, what would stop anyone from putting up these icons? Who would check that they actually implemented it?
Even if that was solved, people would just implement this one thing because it looked good. But there are plenty of other ways to ruin your password security, so you couldn't really trust them more than you could in the first place. (IMHO, this is a core issue with security standardization)
You're quite right that there'd be nothing stopping people from using this dishonestly, except their consciences and the fact they may have some explaining to do if a dump of MD5s of their passwords was released. That may or may not be enough.
In any case, I'm sure that this industry can do a bit better than it is at the moment. With big breaches of LinkedIn, Last.fm and eHarmony in the last 48 hours, surely something can be done.
> except their consciences
I would add ignorance to that list.
This does not detract from the rest of what you said, of course, and I agree that this wouldn't really be useful.
This kind of escalating competition based purely on computing power indicates to me that the very concept of passwords has probably had its day and we should seriously think of better alternatives.
Passwords are no fun to remember and to keep secure for the users either. Anyone with a reasonably active 'online life' suffers from this.
Maybe this is the real reason why Facebook is doing so well? Only one password to remember.
1. The attacker doesn't need the text the user entered anymore, just the precomputed hash
2. Probably the length and alphabet is fixed now, which might obfuscate/protect 'password' or 'test', but reduces the value of a strong password. Granted, this last part is a gut feeling.
But it's still not a real problem since 128 bits of entropy is unguessable in the lifetime of the universe (checking 2^64 hashes a second, which is obscenely many – perhaps every processor on the planet dedicated to the task would be enough – covers 5% of the search space in 34 billion years.)
Is the first case (judging normal passwords) factoring in that a password varies in length? I mean, stupid thought again: You need to test all one character passwords, all two character password, a hash is fixed in its length?
And I wouldn't want to find the original input, I'd want to get in. For that my totally fallible gut says that I'd need to create a 'word list' of hexadecimal character permutations of length x. Is this really an impossible task?
A hash is fixed, but at a long length. Now, because of geometric growth, the shorter lengths are basically irrelevant (since there are 10s times more 19 character passwords than 18 character ones, 100s or 1000s times more 19 than 17, and so on)
On the second point, yes, exactly, you need a word list of all hexadecimal strings of length x. Again, in the case of MD5 (128 bits), this is all the 32 character hexadecimal strings (since 32 characters * 4 bits per hexadecimal character is 128 bits). Such a list has a length of 2 to the power of 128 by definition - 340282366920938463463374607431768211456 items (about 10^38).
Making a list 10^38 items long is not impossible since that's well below the number of atoms in the earth (about 10^50). It is probably impractical however. Suppose you could store the numbers in iron (the most abundant element), you'd need to store each item of the list in about 0.01 nanograms.
It's a perfectly valid post. No one says "use this, this is awesome and secure". If you think the idea is bad, then answer with an explanation. The -1 is simply not useful here. (Neither are one-liners that boil down to -1.)
for (i = 0; i < 1000; i++)
scrambled_password = HASH(scrambled_password)
Aren't we weakening the hash function? Presumably the hash function is not one-to-one, so if you iterate this for many iterations there is a danger that you could end up with a function that has a much higher probability of collisions?Persumably there should be no real reason why HASH(8_char_password) = 160_bit_hash should be less strong than HASH(160_bit_hash).
Not only that, but most hashing algorithms already do several iterations before returning the hash.
1. http://crypto.stackexchange.com/questions/135/why-does-pbkdf...
This is of course not run-time configurable to increase the computational complexity of the password scrambling, but besides that, what are the problems? (I assume that there must be some, since I haven't ever heard of anybody handling passwords this way.)
It sounds good but the challenge, as always, is the infrastructure. I think it would be great if I had a single personal private key from which I could issue chained keys for each domain where I have an account. But imagine managing this across desktops, browsers, phones, game systems, etc. ...
So, right, I was a web developer pushing my PHP-based company to have a more robust-against-db-compromise password hashing strategy. You know what the huge problem was? The huge problem was, MySQL (and hence phpMyAdmin) didn't have a SHA2() function until mid-2010. Not only is SHA2() 'not enough', i.e. it's too fast and you want to do key stretching -- but even then, they didn't even have that.
So suppose you are developing an agile product, someone loses access to their account and asks for a new password, you type `head -c 9 /dev/urandom | base64` into your shell and get back `pYG3fvp9c06m`. If you don't have anything better built yet, you're going to go into the database and write the one-off query `UPDATE users SET pw_hash=SHA1('pYG3fvp9c06m') WHERE username = 'bob.bobertson'`, or, at best, `SET salt='tyDvBBHioUNS', pw_hash=SHA1('tyDvBBHioUNSpYG3fvp9c06m')`.
If you could get an interoperable PBKDF2 working in MySQL/Postgres, PHP, et cetera, devs would use that. It's precisely because it's not easy that it's not adopted.
EDIT: My apologies to Poul-Henning Kamp for implying that he was a journalist. I thought that would be a sort of compliment but I can see now that it's more of a sort of category error. (But I still think that the problem is precisely that the whole dev stack doesn't support any standard.)
He is allowed to say stuff like "But we have yet to find out why nobody objected to them protecting 150+ million user passwords with 1970s methods."
And this is Linkedin. They should know and do better.
I actually imagine that their very gifted developers are running around wondering how they themselves didn't audit this.
or perhaps its that some 3rd party can authenticate users using sha1 passwords i.e. that internally linkedin passwords are scrypted or something, but this dump was from MitM between 3rd party plugin and linkedin?
You can only imagine how many times someone noticed that passwords weren't salted (by comparing stored passwords to a leaked set of hashes or raibow tables after another announcement from some company being hakced) and complained, and got brushed off.
Poul-Henning Kamp (http://en.wikipedia.org/wiki/Poul-Henning_Kamp) is not a "tech journalist."
Exactly how much easier does it need to get? Shall we print out the manual page and put it under people's doorsteps?
It would take like a maximum of twenty minutes for anyone at all, armed with Google and Stack Overflow, to go from "I know nothing at all about password hashing" to "I am securely hashing my passwords" in PHP or any other language. I think it's fair to wonder what the fuck is wrong when, in companies full of tens or hundreds of presumed-competent programmers, nobody does that, ever.
Salting became best-practice in the 1980ies, but the "lost generation" of dot-com wizards never bothered reading "all that old stuff", so they are doomed to repeat the mistakes.
And I think LinkedIn really has no excuse.
It turns out that this is a bit complicated, as my colleagues readily pointed out to me. So for example you are basically saying "write a .PHP file and execute it locally," which is perfectly fine, as long as the problem comes from your boss or one of your testers -- it is risky when it occurs on a production server (because the script you're generating is insecure). On your production server you really do want to execute the action from within a MySQL prompt if it's possible, and so it's a sort of two-sided game of "I'm going to reset my password over here and then update their (salt, password) with the result of my local PHP queries," and that's a bit weird as a process.
The other tender point is that once you've made a choice, it's very hard to change it. So, "all of our existing passwords use the old system, we're not changing!" was a very strong argument and I did have to spend a bunch of time creating a fall-back for legacy passwords.
I would agree, however with this: in general there is a reasonable expectation of, "if we're doing this so much that it bogs us down, then the app is mature enough for a proper email-sending password-reset tool; and as long as it doesn't bog us down we'll do it the hard way." But convincing people to make the hard way even harder is a tricky proposal even on a good day. It's like telling people, "no, leave that code, I know it does 2^n operations but n is always small and it's not actually the part that slows down our system and it's more readable this way." The intangible -- security/readability -- is being negotiated for the tangible -- dev-ease/speed. I had trouble selling it to the other devs.
I don't think I ever want to be _that_ agile. My agile projects usually have a set of application functions exposed as scripts immediately. And yes, proper password change is one of them. (besides, how about just using `pwgen 16` and not some trickery with head and random?)
Second goal: establish a process that gets everyone flagged that tries to change things using phpMyAdmin that have proper equivalents in your scripting toolkit. Agility is no excuse for sloppiness. If the agile crowd still insists to be agile to death, call the whole thing MVT (Minimum viable toolkit).
Using a framework where all this can be done from a REPL also helps a lot.
But no, I don't think it is at all obvious why LinkedIn used unsalted SHA1.
LinkedIn went through an IPO, which implies that a number of companies have audited them from head to tail several times along the way.
If the commodity you buy is millions of user accounts, shouldn't you, as investor, at least check that there was a lock on the door to the warehouse ?
I have a question though, how would a strong password that takes around 1 second to hash affect the scalability of these systems; would it impact the login times of users a lot? Imagine thousands of people trying to login at once. Might it be the reason linkedin didn't hash and salt properly?
2. Any advice on encrypting passwords? We store passwords for some 3rd party services for our users.