One way to fix your rubbish password database
blog.jgc.org
blog.jgc.org
http://akibjorklund.com/2009/supergenpass-is-not-that-secure
Let's say that "foo" and "bar" are two distinct passwords that have the same MD5 hash. Then bcrypt(md5("foo")) == bcrypt(md5("bar")), regardless of how bcrypt("foo") compares to bcrypt("bar"). By pre-hashing with MD5, you have added possible collisions that weren't there previously, and those collisions remain regardless of how many more hashes you pile on top.
It seems to me that the most obvious problem is that you get two chances at colliding — once with MD5 and once with bcrypt. But bcrypt is not known to be especially vulnerable to collision attacks, so this setup is probably not noticeably worse than MD5 alone. But that's just looking at probabilities — I ain't no fancy crypto expert or nothin', so there might be much more subtle vulnerabilities than the added chance of collision.
Yeah, that's all I'm saying. I was answering a question about being "fundamental worse," and fundamentally, there are now two sources of potential collisions instead of one. In theory, that's twice as insecure! However, the practical effect is unlikely to rise above absolute nil anytime soon.
As a practical matter, we can basically say that never happens. Certainly not for passwords that are user selected and not designed to collide. And since the hash itself is hidden by bcrypt, the attacker won't know md5("foo") even if they were inclined to find a "bar" with the same hash.
Cryptographic pseudorandom number generators also collide and produce cycles. MD5 hashes of arbitrary ASCII strings are reasonably modeled as random numbers, and the concern you cite is unmeaningful.
An upper limit on string length suddenly makes some sense.
The idea is that you have two messages, m1 and m2, or m1 and m1' if you prefer, and you vary bits in both until you get a collision. You need some area of m1 and m2 that doesn't matter for the application, so that you can change those bits and find a collision. Since all bits of m1 are supplied when entering the password, you have no ability to modify it without getting the user to change his/her password.
If you could collide any arbitrary m1 as it's given to you, then attacks like fake certs with signed MD5 hashes could create the fake cert after submitting it to the CA and getting the signed cert back, rather than before.
Also, the collision process requires knowledge of m1 so you can see the intermediate hash states. If you know m1, the password/passphrase, why are you trying to find a new m1' that hashes to the same value rather than using the pass you already know?
An attack of concern for using MD5 as a password hashing step would be a first preimage attack. [2]
[1] https://www.google.com/search?q=md5+collision+block+birthday... (first link at present is http://www.win.tue.nl/hashclash/SingleBlock/ )
If it were me doing it, I'd take the article's approach as the first whack, and then as each user logs in, validate their password (as per the article), and then also set their password to scrypt('salt', 'password') vs. scrypt('salt', md5('password')), so that they were current. Then just set a flag on the user record like "new_password=True" or something.
That gets you the stopgap without having to muck around with scrypted hashes forever. Send out a few emails to your user population with a note that you've upgraded your password strategy and that they should log in. You're still not going to get 100% coverage, but at least for whoever you don't get to log in, their passwords aren't 'in the wind', so to speak.
Edit: I'm not sure if you edited yours, or if I just read it poorly, but I think we're saying the same thing. Note, the article's strategy is basically your first scenario -- just bcrypting (or scrypting in the article) the existing md5 hash.
I added a second password column to the users table, then ran a script that queried the table for the existing hash, generated a bcrypt hash from that value and wrote it into the new password column. Then I removed the old password column. No need to wait for the user to log in.
When people log in today, the code takes their password, runs it through the old hash routine, runs the output through the new hash routine, and compares that to the password on file.
I think a better idea would be to establish an easily implemented pattern for "password bankruptcy" that companies could follow in the case of a leak.
One thought is to invalidate all passwords and fall back on email password recovery when a login is attempted.
This leads me to an idea I've tried once - if access to the inbox is equivalent to password credentials, why not use an email to login? By this I mean the web site login is a single field - email address. The system emails a one-click-login URL to the user that can be re-used (possibly with a month expiration time). The user can look up the URL in their inbox when they want to login again, or use a long-lived cookie.
In practice I end up doing this for little used sites because I use either my phone, tablet, and two laptops for browsing the internet.
It's annoying if you work somewhere that doesn't allow access to personal email accounts and you want to log-in to something.
Couldn't you just run your whole database through X more rounds of MD5 and do the same in your authentication function?
That way, script kiddies couldn't use precomputed rainbow tables they downloaded somewhere off Bittorrent.
Each additional round will also reduce the speed of a brute force attack while still keeping the changes to the codebase will be pretty small.
Unless there are rainbow tables for a certain number of MD5 iterations, it would be a start...
Are there any actual arguments against using this as an 'easy' fix to the precomputed rainbow tables scenario? Multiple rounds of a cipher seem to be a relatively common operation in crypto and have helped other old ciphers. One of the more prominent ones would probably be the move from DES to triple DES.
I guess dictionary attacks on GPUs would still be easy enough, even with more iterations, but anything that isn't directly in a dictionary might benefit quite a bit from multiple iterations.
It's not as good as actually using proper crypto rather than hashing algorithms that were designed to be fast, but it seems like an easy to implement low-risk solution.
The main advantage is that it would still keep people from using precomputed rainbow tables and slow down brute force attacks with a minimum of additional code, wouldn't it? (similar to the switch from DES to triple DES back in the day)
Using a fast hash algorithm for storing passwords is fucking braindead and a DOA decision to make about security.
Your "solution" doesn't solve anything.
Your system would seem to be practical if you know there are no weak passwords or if you dont care if only some of the accounts are compromised.
You've also got to watch you don't DoS yourself.
In step 4, we make the assumption that their API is out in the wild, in use, and sends the md5(s, p) in the request. I get that we take that value, run it through scrypt and match against our stored value to authenticate. So the database has:
scrypt(s', md5(s, p))
No problem authenticating the API requests with that.Step 5 says once the user logs in with their actual password, we update entirely to the new scheme of scrypt(s'', p) and store just that. Now the database only has:
scrypt(s'', p)
But the API user still sends md5(s, p) to authenticate, right?So then what happens when that same user goes back to the API-using app? It's still uses the API so it'll send the MD5(s, p) and fail since we've discarded the transitional scrypt value when they logged in via the web interface.
Is there a deprecation period that supports both types while API using apps updated to a new API for the new scheme?
If your original database contains a bunch of unsalted SHA1 (or worse, MD5) hashes, what good does securing the hashes themselves do if the means to generate the corresponding plaintext has already been released into the wild?
Someone please tell me I'm missing something obvious.
If my password is "password", and I change it to "#b1@password%3dy", and then hash it, isn't it secure from basic dictionary/rainbow table attacks?
I'm a bit new to cryptography, so please forgive me if I'm not understanding some of this correctly.
Therefore, if you are storing the per-user salt as the first bytes in the hashed password field, then you have to be careful when you "throw away the old weak hash hi and forget it ever existed."
It works well.
... then you should stop doing that and you should start using OAuth, so the client application never sees your user's password.
[1] http://www.robertsradio.co.uk/Products/Internet_radios/STREA...
Me mentioning it: http://news.ycombinator.com/item?id=596126
The attack: http://news.ycombinator.com/item?id=639976
yes.