Md5crypt is no longer strong enough
phk.freebsd.dk
phk.freebsd.dk
If your password database is leaked, there are 2 categories of attacks that you need to be concerned with:
1. Brute force
2. A weakness in your implementation
He's right that brute force attacks will require potentially more work if you have a custom scheme as the attacker will have to work out the details of the scheme before using brute force (category 1).
However, many cryptographic compromises are actually due to bugs or weaknesses in the implementation (category 2); i.e. on paper your custom combination of schemes is fine, but you made a bug in implementing it or forgot to handle some case correctly.
How can you have confidence that your scheme is safe from category 2 attacks? Use a scheme that has already been widely reviewed and attacked. If no one has been able to find a vulnerability in it thus far, there's a good chance that your attacker won't either.
So there _is_ an advantage to using the same scheme that many others use. Use something hard, like scrypt, bcrypt, or PBKDF2, but also find an implementation of it that is likely to have been well used, reviewed, and hopefully attacked.
It will not deter targeted attacks (spearfishing) by competent attackers, but again, not much would.
Incidentally, what kind of drive-by attack are you thinking? Like, you have your password database sitting out in the open and there are bots crawling the web trying mechanically to crack it? Because that doesn't sound like a common case, but I'm not sure what else you could mean.
This comment was factually incorrect and only served to cloud the issue. The link given in one of the responses provides more than enough information for a developer to understand the options available for password storage.
Apologies for the noise.
So no, if your scheme is “simply” bcrypt(password), it will not take so short a time for someone to extract passwords from rainbow tables, because they won't be able to have rainbow tables. Proper rainbow tables would need to both account for the salt and the number of encoding runs your bcrypt function has (which, incidentally, is also a configurable value).
Some information on how bcrypt, scrypt and friends apply salting and stretching to resist attacks was in yesterday's post at http://throwingfire.com/storing-passwords-securely/ . It's worth a read.
Thank you for clarifying!
All major internet sites, anybody with more than 50.000 passwords, should design or configure a unique algorithm (consisting of course of standard one-way hash functions like SHA2 etc)
So he is in fact telling you to use a standard one-way hash function along with whatever custom implementation.
For example (while collision resistance is not that important for password hashing), by chaining MD5(SHA1(x)) you've created a better chance of collision than if you just used SHA1(x).
If you use 10 hash functions in any possible combination, you must evaluate the security of every possible combination, or prove that your resulting function is secure with any combination of these 10 functions. And for what? For only a slightly larger, constant, cost of attack, which you can easily get by adjusting a parameter in scrypt.
(Note that even using a standard block cipher requires a separate evaluation for different purposes -- you can't claim that your hash function/MAC/mode of operation is as secure as AES only because it's based on AES.)
If you read my replies in this thread, I emphasize the point that if you really need a custom function with secret parameters, use a single algorithm with a secret key (I even gave a link to one such scheme with HMAC.)
If 'key' must be kept secret in order to achieve its stated security properties now the system is much more unwieldy for the defender. He now has a key distribution problem that he didn't have before. If the defender can reliably solve the key distribution problem then the attacker wouldn't have gotten a copy of the password key database.
On the other hand, if 'key' is public, then the incremental computational cost added with AES raises the cost for the defender, but not the attacker. The attacker can just un-AES a password hash once and remember the result whereas the defender must perform the AES on every password validation.
Many of the attacks that lead to password hash exposure involve a SQL injection that dumps the password table. If you store the password in the source code an attacker must then capture both the source code and the hashes. Furthermore, if they leak your key, you learn that they have gained access to your source code.
I agree that this is a less than ideal system. My point was not to design a better system, but to point out that composing permutations with a secure hash function is at least as strong as that secure hash function. If you disagree consider the identity permutation. Really you'd want to implement it like SHA256(AES(key,SHA256(password|salt)).
If I wanted to design a better system this is what I would do. Put the hashing system on a dedicated box that maps passwords to hashes. To compromise the hash key an attacker must compromise your hashing box and your database. Locking-down a box that performs only one action, hashing, is significantly easier than locking down a complex web application.
This is the part that I think is subtly invalid. Coincidentally, I was having this mailing list discussion the other day: http://lists.randombit.net/pipermail/cryptography/2012-May/0...
It has to do with the fact that the defender is obligated to pay the cost of the AES every time (because he follows the rules of the algorithm by definition). But the attacker is free to bend the rules however he can and one thing he can do is perform the AES once on the entire data file, a negligible amount of computation.
So even though AES is an identity transformation, it makes the function more costly for the defender than for the attacker. Wncrypting one's database may be an acceptable tradeoff for other reasons, but when we're considering the work factor for cracking passwords this is worse.
Think about if instead of just one iteration of AES it were iterated a zillion times in order to make it expensive. That would be a complete failure.
I agree, and I don't think AES (or any easily invertible function) should be used to increase the work factor.
We have been arguing past each other. I was responding to a post that said ~'if you compose two functions and the second function is insecure, than the composition is insecure'. While such a statement is correct for compression/hash functions, it does not hold for permutations.
Do you agree that keys or peppers decrease the risk of passwords being revealed if the key or pepper can remain secret from the attacker? Assuming that keeping a key secret is much easier than keeping a bunch of hashes secret, which is a fair assumption, shouldn't such a method increase the security of the system?
I'm not familiar with this term 'pepper'. I take it to be a key with which you encrypt a password has as in your construction?
> Assuming that keeping a key secret is much easier than keeping a bunch of hashes secret, which is a fair assumption,
I don't know that it's such a fair assumption. For example, where does the server software get the key when it boots? From the hard drive? Is it something that is chosen or typed in by admins? If so, how do you ensure this key isn't itself simply derived from a weak password? If it's a password, how likely is it to be the same as the root login password or master database password? If the app developer re-uses this key or password in multiple contexts it's an additional exposure.
> shouldn't such a method increase the security of the system?
I think it could help and might even be a good idea overall.
That said, I don't think we should consider it an intrinsic part of password hashing security because:
a) It might help some with pure SQL injection exploit scenarios but anything beyond that it probably won't help at all. So any security benefit it brings is very hard to quantify.
b) It could apply equally to any sensitive database field that didn't require join capabilities.
c) It imposes a work factor to the defender which can be bypassed by the attacker.
I think the most secure method would be to store a random string in the persistant memory of a hardware dedicated device that performs the encryption on your behalf (like a smart card or trusted computing module). If that is not possible you could still put the key inside the hash program and give the app only execute privileges for that program (the key is safe unless the attackers get root).
a). If it protects against a very common attack or attacker why not do it. Security should be layered.
b). agreed
c). The work factor of a single AES encryption is negligible for both the attacker and the defender and can be ignored. Furthermore it can be removed entirely if you hash both the input and the output as I proposed in a previous message, SHA( AES(K, SHA(password) ) ). Now the attacker has to run AES for each try as well (but it doesn't matter since any reasonable work factor will be orders of magnitude larger than a single AES operation).
I don't think that's a good idea either. There's still a ton that can go wrong that won't be apparent at all to someone who knows how to analyze entropy flows through the internals of hash functions. There's a reason why we have a PBKDF2 that's favored over the original PBKDF.
He probably means to use some combination of well known/tested algorithms rather than inventing your own crypto, but I think his wording is ambiguous enough to be dangerous. While there is some benefit to using a unique algorithm for your site, it's almost certainly more risky than using a secure algorithm (i.e. bcrypt/scrypt/etc) even if every other site was using it too.
But honestly, I'd be pretty happy if we could get all sites with 50,000+ users to salt and hash their passwords with any algorithm.
Here's a question for the people here who actually know wtf they're talking about: if I choose to iterate through a set of hash functions with each pass of PBKDF2 rather than using the same one each time, what effect does that have on the entropy of the system and so on? Would it make it easier to crack, or harder?
This whole line of thought is at best a waste of time, and at worst dangerous.
PBKDF2-HMAC-SHA-256 is a vetted NIST approved standard with an adjustable work factor. It has been subject to professional attention for many years.
BCrypt, while not subject to nearly as much analysis nor approved by NIST, was designed by very competent cryptographers, has an adjustable work factor and is based on blowfish -- which has been subject to substantial professional attention.
Use one of the above with the highest work factor you have the processing power for and call it a day. Don't try to roll your solution.
I'm far too lazy to do anything other than slap bcrypt on it, unless there's a pressing need to do something else, which there never is.
Hash functions have different requirements than key stretching functions, but if you're interested in the security of combining cryptographic hashes, Google for "combining hash functions" and "chaining hash functions" -- there's a lot of interesting research.
Yeah, that now seems obvious. The input string is the same, so the information entropy is the same. I'm struggling to think of the concept that I need here, I want to say kolmogorov complexity but I know that's wrong too.
NO! Bcrypt in particular is designed to resist GPU brute forcing.
If you're being attacked by anyone other than opportunists, you have bigger problems than your hash function. As soon as someone attacks you specifically, you're in a "trust no-one" situation, and suddenly it's time for anonymous meetings in basement carparks and the like.
If the answer is "my hashes protect something that is particularly valuable," then the attacker probably isn't going to hack your hash function, he's going to hack your secretary or your garbage disposal or something like that which is more effective.
Of course, in practice you should just use bcrypt anyway.
I would argue that almost everyone that is storing passwords should start worrying about people bring racks of GPUs to bear against you, because it is so cheap. At 33.1 Billion MD5 hashes/s with 4 dual-linked GPUs (one machine), you can eat through all 8-digit alphanumerics very quickly for a few thousand dollars. (of course that is using PBKDF1 or less). I had done the calculations in a spreadsheet and forget how long it would take, but it is way shorter than you'd thik.
Tie a card-carrying cryptographer to a chair until he delivers it ?
Alternatively, implement it very, very simply and put it on github, and ask card-carrying cryptographers to vet it. That one is harder on the ego though.
Also, many big service providers still use MD5 or SHA1. I saw a company migrate to Google Mail from an old legacy mail system and part of that was integrating some authentication services too. This was several years ago, but at that time, Google accepted hashes in two formats only... MD5 or SHA1.
if( user has an old style hash ){
if( password verifies against old style hash ){
add a new style hash
delete the old style hash
log them in
}
} else {
if( password verifies against new style hash ){
log them in
}
}The issue is you don't always have the ability to do the old style, that's why people prefer a hash that they can be assured will be found everywhere.
Edit: there's more detail at the link below. It looks clear that in at least some of their schemes they deliberately do not send the client password to the server, which sounds like a decent idea.
http://www.skullsecurity.org/blog/2012/battle-net-authentica...
I wonder, though. Could there be a "code has changed" warning from the client? I mean, authentication should be pretty damn stable, and maybe even universal. If someone does modify the page, it'd be nice to know if that change was reflected on other sites, and it'd be nice to know that someone I trust had signed off on it (cryptographically).
A simple alternative is to build it into browsers. A password field could generate a salt per-domain and automatically encrypt any queries to password-fields. The server doesn't even need to know about it. You'd have to be more than a little careful building it, obviously, and you'd have to find a way to deal with passwords used on more than one site, but it could work.
How do you degrade for NoScript users?
And where's the harm in that anyway (assuming TLS)?
Of course, if you're salting the hash uniquely for each user, then this approach isn't very helpful.
Finally, if your server will accept hashed passwords, then getting the hash is as good as getting the password for access to your site. The only benefit is that the password is hidden so you may avoid compromising security on other sites.
I've seen this advised a lot. However, where are you going to put the per-user salt? I presume in your users table, or somewhere in your db. Does this not mean that the salts are just as likely to be hacked as the encrypted passwords?
Are you not better off with either a single salt that's stored somewhere outside your db or with some other scheme for algorithmically picking a salt based on the user id?
If you store straight unsalted hashes, you can precompute the hash for every likely password, store the precomputed hash -> password mappings in a nice efficient data structure (http://en.wikipedia.org/wiki/Rainbow_table) and use the same precomputed tables to reverse every hash in the system with a very fast lookup. And every other system using the same unsalted hash scheme.
If you have a salt, this doesn't work -- you'd need a set of precomputed tables for each salt value, which rather defeats the object of precomputation.
If you take it a step further, and use a storage scheme like bcrypt or PBKDF2, not only are you protected from the precomputation attacks, but testing each password candidate takes much longer than a straight cryptographic hash -- so brute force attacks becomes much slower, too.
With salts, the attacker can only attack a single user at a time. Try a password with one user's salt, see if the hash is in the database, try the next user's salt, see if that hash is in the database, repeat.
Salts don't need to be secret to work.
The best approach to securing information going from client to server is SSL.
How so? Surely the client would know their username too?
I agree that for small sites, changing the auth check and updating all the existing rows (and deleting/overwriting the old hashes) is probably the best solution though.
No, just SHA1 the password and check, and then bcrypt(SHA1) and check, until the encryption process finishes. Or just check the length, or store the hash type along with the hash, a la Django. These problems are trivial, really.
I'm sure there is a way to improve it, but I don't know how, sorry!
1. bcrypt(SHA1(pass)) right now to secure all pws
2. check against that, then update to bcrypt(pass) on loginAFAIK its this
1. Continue using legacy, less secure hashing algorithm
2. Upgrade your password storage scheme and carry around some extra code and/or fields in your db.
3. Crack you own users passwords and convert them
Its all kind of smelly but 2 seems like the only option (unless I'm missing an option?)
Regular users should have no problems if there's an accompanying blog post Non-regular users might just remember about your site/service and come back :)
5. Use bcrypt/password stretching. Store the work value alongside the password and upgrade it as people log in. To me that's not really keeping legacy code around; just an extra variable...
That's, ummm, _concerning_…
I don't suppose anyone knows how securely Google are storing my gmail password? I _hope_ it's not unsalted MD5 or SHA1. (Especially since google search is probably the best general purpose md5 hash reverser most people have access to.)
But you're right, as far as anybody but you and Google are concerned, it's encrypted.
Edit: never mind the last parenthetical; it pretty much wouldn't help in a database leak at all (just adds one extra hashing step to the cracking process), sorry. Still helps for non-HTTPS logins though.
Edit 2: ... though maybe if the client-side hash were something strong like bcrypt, it would help in the case of database leaks on HTTPS sites that refuse to use strong hashing on the server side for performance reasons. Sorry for the rambly disorganized post.
This is not theoretical. This is what Tunisia did to Facebook and it's what online banking trojans (e.g. Zeus) do every day.
It's not surprising that they don't support importing custom formats. They support two formats to bulk-import and then surely convert to their standard from there.
The attacker would have to recompute all hashes for each user using their individual salt.
At least that's what I remember from Computer Security ha
No, I'm obviously not proposing that people do something stupid with crypto, we've had enough of that in recent days already.
But I am trying to provoke one or more card-carrying cryptographers to realize, that while password protection may be a problem we have good and strong theoretical solution for, those solutions will not protect any passwords until somebody turn them into Open Source code we can use.
I only wrote md5crypt because nobody else had done so, and nobody else wanted to do so at the time, and FreeBSD needed an ITAR exportable password scrambler.
If more cryptographers wrote more code under liberal Open Source licenses, instead of bitchy complaints against the people who do write code, then the world might gradually become a better place
I have been dreading this announcement for a couple of years, knowing full well that the majority of the world can't tell MD5 from md5crypt.
It is a credit to hackernews that you could, much appreciated.
If you really want a "custom algorithm", the better advice would be to use a known algorithm with a secret key.
For example,
https://wiki.mozilla.org/WebAppSec/Secure_Coding_Guidelines#...
Why? See http://en.wikipedia.org/wiki/Kerckhoffs_principle
In most case, you don't need this, though. Can you prove that your construction is secure? How many hashing algorithms are there? Remember that big O thing. Just increasing the workfactor for scrypt will be much better than using a combination of 5 or so hash functions.
This line of reasoning assumes A) that the probabilities of attacks on the algorithms are independent, and B) that the algorithms in use do not substantially reduce their input entropy, both of which are potential attacks.
It does _NOT_, however, assume that the scheme be kept secret.
It's a pity that seemingly trivial changes on an crypto algorithm can totally obliterate its security.
The app provides a fixed 'secret key' that defines (somehow) a sequence of hashing and scrambling functions to apply to the password. So, for example, the secret key "df8dfuhuejew3" (or whatever) might be interpreted as '3 iterations of md5 hash followed by rotate hash 5 places to the left followed by 2 iterations of bcypt etc..'.
So, in effect, each app (with a different secret key) will have a different hashing process, but with the advantages that its repeatable and cross platform (reference implementations could be available in the major languages).
Of course, if your secret key is hacked you are in trouble, but at that point the hacker can probably get your source code for whatever hash system you use, and anyway, the combined hashing process is probably slow enough to prevent brute force attack.
Any thoughts?
Edit: Probably better to just use bcrypt with a larger work factor.
bcrypt(password + hardcoded application salt + db-stored user salt) is far less complex and arguably just as secure, since your argument is predicated on "assume the source code is safe". bcrypt is tunable to be as slow as you want, even to account for hardware progress.
I did not know that. Can you expand on this point?
As hardware improves, you just implement a system wherein when the user submits login information, you verify their password with your old work factor, and if it passes, re-hash with your new (slower) work factor and store the updated hash. This allows you to effectively use a progressively slower algorithm over the lifetime of your application to compensate for Moore's Law.
You obviously don't want to pick a work factor that's too high for your web server hardware, since that opens you up to DOS attacks, but a reasonable work factor can easily mitigate the weaknesses of MD5 and SHA1 - notably, that they can be computed by the hundreds of millions of second on the right harware.
If anyone else is wondering how to implement bcrypt with a cost parameter in PHP, see http://php.net/manual/en/function.crypt.php under CRYPT_BLOWFISH.
That the security of a cipher system should depend on the key and not the algorithm has become a truism in the computer era, and this one is the best-remembered of Kerckhoffs's dicta. ... Unlike a key, an algorithm can be studied and analyzed by experts to determine if it is likely to be secure. An algorithm that you have invented yourself and kept secret has not had the opportunity for such review.
http://en.wikipedia.org/wiki/Kerckhoffs%27s_principle#Implic...
Moreover, you made the security of your system "key"-dependent: what if I generate such "key" that will only use 5 iterations of MD5 and 1 iteration of SHA-1? This would be a major failure. Imagine if the security of AES was not 2^128, but varied between 2^10 to 2^128 depending on what key you supplied -- would you use it?
Agreed. The unpredictability of the work factor would be a problem.
By that time we'll all know that your password is "I Love Care Bears 123" :P
md5crypt is a "poor man's" crypto system, simply because it's built on MD5, which is intended to verify data integrity more than it is designed to hide information. We've known that md5 was broken since 2004, so at this point, anyone actually using MD5-based systems for password hiding really has no excuse.
Quantum computing is a whole 'nother ballgame, and while it does offer the promise of making current encryption schemes more or less obsolete, it's an entirely different ballgame than "chain together a few Radeon cards and crack a DB full of MD5 hashes".
Keep sensitive information on your own storage media, not on some website's servers connected to the internet.
(pause for collective gasp from the nerds)
Until it becomes a crime of some sort to use these simple, well-tested hashes to store passwords and the world police begin locking-up CIOs because of it, then don't expect anything to change.
I'm just surprised to read about its deficiencies on Hacker News as if it were news.