How to Safely Store Your Users' Passwords in 2016
paragonie.com
paragonie.com
The operator on the line asked me a ton of questions starting from my user-id (but not password), date of birth, full name, father's name, place of birth, type of account and many others that I don't remember now. Only after I correctly answered all these questions, did he start acting on my instructions.
But if i tell them to not make it that easy, i cant manage my shit over the phone anymore :/
Seems that they don't have any other way to grant the technicians access to the server/my account, which is absolutely ridiculous.
We live in an age where this should be unacceptable. Why aren't there financial security laws yet?
I don't think banks actually care about individual user login security too much. Credit Cards reallllly suffer from security breaches though
But the only reason that banks care so much about credit-card security breaches is that the law forces them to do so. If the law didn't make credit card fraud the bank's responsibility, then they'd be just as lackluster about preventing it as they currently are about securing login credentials.
Remember when you couldn't take your cell phone number with you and so pretty much nobody switched carriers? It was a massive pain. Now it's easier than ever to switch, except most people are locked into multi-year contracts. Switching friction = high, but not impossible. As you said, TMO is trying to compete here.
Cable has monopolies on towns, so there's 0 incentive. People couldn't switch even if they wanted to. I suppose there's satellite, but you'll still be paying the cable company for internet -- they get their pound of flesh no matter what. Switching friction = impossible.
https://paragonie.com/blog/2015/08/you-wouldnt-base64-a-pass...
Furthermore, credit unions can't make risky bets that could put them under. The money you deposit goes out as loans to other people.
I avoid banks like the plague. Too shady for me, never again.
Where did you spend your honeymoon? What was the name of your first pet? What is the name of the street where you grew up?
For any given person, a LOT of people know the answer to these kind of questions.
Also, I hate it when people use date of birth to verify identity. Medical people love doing this. Um, just check the person's Facebook and see when everyone wishes them a happy birthday, then go access their medical records?
answer = PBKDF2(hmacsha1, password + question, "", 100000, 16)
This is also incidentally the basis for how I generate unique passwords for every service except banks, communication, and other sensitive things. I want a different password on every website and don't want to trust any password-remembering software I didn't write. The same function works fine for generating answers to secret questions.By that, I mean the overall security of your password scheme is analogous to what people get out of a password manager.
Still, a lot better than password re-use.
KeePass? It's great, and open source.
It's not quite as bad as asking "what species was your first pet?" but not much better.
This is the real problem with security questions; the answer space is often so very small, and can be narrowed down even further with a little research.
Questions like "what was the first name of (your maternal grandmother, your first best friend, etc.)" are very common -- well, there are stats on most popular first names of given generations in different places. If you know what country the person is in, you can make a good guess at these.
Tangential to the actual issue, but in that field it's to prevent patient mixups, not to defend against malicious attackers.
As you implicitly point out, however, it doesn't require any portion of the password ever to be visible to the call-centre employee; one can just supplement an individual hash by a collection of hashes of appropriate character subsets, and then (say) randomly pick among the available subsets.
I went to my boss and explained that we can't do that. It's inviting exploitation. He responded to me that we had to keep them in plain text, in the database so that we could send them to users who forgot. If they can't login, they won't order product.
I have heard similar stories from other IT professionals. It's amazing that these operations aren't getting pwn3d twice a week.
bcrypt.genSalt(10, function(err, salt) {
if (err) return; //handle error
bcrypt.hash(clearPassword, salt, function(err, hash) {
if (err) return; //handle error
// Store hash in your password DB.
});
});edit: Looks like it depends on which version you use.
'npm install bcrypt'[1] gives a version which supports true async usage (via V8 async callbacks in native code)
'npm install bcrypt-nodejs'[2] gives a pure JS version which I linked to. Which is the top search result for 'nodejs bcrypt'
[0] - https://github.com/shaneGirish/bcrypt-nodejs/blob/master/bCr...
[1] - https://www.npmjs.com/package/bcrypt
Kidding.
... and conversely a 1 second KDF in Javascript will be trivial to crack brute force. There is no point in it.
Please don’t pull numbers out of nowhere.
edit: and I'd assume a GPU version would be faster.....
bcrypt is used on the server (node.js) to hash a user's password before storing it in the database. Then later, when a login is done, the password from the login is checked against the stored hash to see if the match (and by implication, that the original passwords are the same). So the usage is mostly on the server as expecting a client such as a mobile device to do large rounds of CPU intensive work just to login isn't going to make for happy customers.
So to me the logic behind a pure js version is one that can run on the client too. It actually sounds interesting, as it could improve password security! But it's just too slow.
In my case, I ended up sha512'ing the password client-side, and use that as the "password" sent to the server. If there were a native KDF for all browser's i'd use that instead.
Some windows developers have issues compiling native modules, and it was easier to try the native plugin, if it fails fallback to the pure-js version with an identical API during development. That way we get the ease of development with pure-js, and the production speed of the native module.
Also there is the ability to run client side (mesh networks with webRTC), and the ability to run this on architectures other than x86 where the native version won't compile.
Important to note that the only advantage of hashing or prehashing passwords on the client is offloading work from the server. It doesn’t improve security (on a website).
Re: GPUs:
- One of bcrypt’s advantages is that its memory requirements make using a GPU provide much smaller returns compared to, say, iterated MD5: you can’t parallelize it as much.
- The point of using GPUs to generate password hashes is generally only to break them, because GPUs can do lots of hashes in parallel. Your web service probably doesn’t need to generate that many hashes at the same time, and it probably isn’t otherwise running on a server with a computey GPU. Clients’ GPUs? Even less parallel, they’d generate on the order of one hash per year.
There, the encryption is handled asynchronously, the event loop is yielded while the calculations take place in a C extension.
Unfortunately a lot of people think that these functions are a simple way to "not have to deal with callbacks", which is not correct...
Is there any reason not to use the Sync functions if you're not writing a web server and don't particularly care for performance?
This may sound like "not a big deal", but even an operation that takes 1ms will cause noticeable lag at ~50 operations per second, and block everything approaching a thousand.
For emphasis: they block the entire event loop. Not a single request, not a single client, not a single operation or event. The entire application for every user everywhere.
How does that look now?
[1] at the bottom of this page: https://medium.com/the-story/signing-in-to-medium-by-email-a...
But as an IT manager, you could enforce people from your company to use their work email address for accounts that are directly linked to your company activity. For instance, We use a online tracker[1] for Use cases management. When someone leave the company, it does not have access to its email account. Consequently, he/she would no longer have access to our tracker by not having access to its email address. This could rather convenient scheme.
[1] Pivotal tracker (who requires password to connect).
In a near future, we could have a "Argh I lost my email address" link replacing the now very common "I lost my password". Nevertheless, I have no clue what would be a standard process to recover a lost account because of a lost email address, maybe a SMS based process and personal questions like "what was your first pet name?"
No it doesn't.
No issues with incognito mode.
"explain why you forgot your password"
I don't know why but that made me laugh.> Email addresses on Medium are case-sensitive.
Wow. Is there ever any possible benefit to this? (Serious question)
In practice, most email providers don't actually honor that, and so in practice it's a bad move. It is technically correct though.
I guess I never thought to look that up. Likely due to the common way being so pervasive.
> The local-part of a mailbox MUST BE treated as case sensitive. [0]
We lower-case all e-mails in our systems. I'll let you know when we hit a system that actually treats it case-sensitive.
"the local-part MUST be interpreted and assigned semantics only by the host specified in the domain part of the address."
- Private key is stored in a secure place. Offline for all i care; printed on a piece of paper; memorized and swallowed.
- When user creates the account - password is padded with salt, then a public key is used to encrypt it. The resulting encrypted form is stored, along with the salt.
- When user attempts to authenticate - the password that is provided is padded with the stored salt, encrypted with the public key and compared to the stored password.
Private key is never used when comparing passwords. Never available to the system doing authentication, etc.
The only purpose of using a reversible encryption - is to be able to switch to a different authentication provider completely transparently to the user.
I've implemented this functionality considering that we may need to switch over to active directory (or some other directory) in place of storing passwords in the database - but never used it, fearful that i would be committing some cardinal crime against proper security practices
Thoughts?
If you're switching to a non-password based system like OAuth, make users link their accounts after logging in with their password. If you are going to completely drop the password, then send each user an identifiable link where they can link their account.
Our system already went through a similar transition years ago - when we had to migrate away from NDS (Novell), since we abandoned that technology.
In NDS we did not have access to hashes. We ended up having to "MTM" the passwords in opur web app to intercept them before they are sent to NDS and after a successful login store them in the database. To transition our entire user base took a very long time (more than a year), since we had no way to compel the users to login and had to wait until they login on their own volition.
Ever since we've been reluctant to lose the ability to recover the passwords. Transition to Active Directory has been on our agenda for quite some time - and having the ability to decrypt the password for this purpose - we can make that happen in short order (AD hashes the passwords, so it will be a one-way operation of course).
But the question still stands - is asymmetric encryption strong and secure enough to take place of hashing, assuming the private key is secured? I have not been able to find any relevant sources to answer these questions.
It should be (though it's a big "assuming"). Being able to quickly move away from an older less secure hashing scheme rather than waiting for users to log in could be a security advantage too.
Interesting point about it being an advantage. I am actually going to make a note of this.
One way of handling the private key that i proposed - was to use a privileged access vault like Thycotic. Or just have it on a USB key in a lockbox.
If you store it on the server alongside the encrypted password and salt, you're effectively using a bigger salt with a weird hashing algorithm that hasn't been vetted by any expert for the specific purpose of password storage.
If you store it on the client side, the user needs to supply the public key whenever he tries to authenticate. This has the effect of making the salt much bigger, but you don't need to use public keys to achieve the same effect. If you can trust the user to store a public key, you can trust him to store any large chunk of random bits.
So apart from the virtually useless fact that you might be able to decrypt the password with the private key (why would you even want to do that, instead of just resetting it?) your method seems to offer no real benefit compared to simply using larger salts with argon2/scrypt/bcrypt/whatever.
We are using RSA. The public key is on the server. We simply using RSA as a "hashing algo". Randomly generated salt is added to the password sent by the user to make the password (much) longer as well as strengthen it before the encryption takes place.
I have explained the reasons for storing the passwords (for the time being, hopefully temporarily) in other replies in this comment branch:
quote:
There's no argument from me regarding undesirability of keeping the passwords. But we also have a different concern - we don't want to be in the business of authentication at all - our goal is to have a third party service or appliance performing authentication and forwarding us already authenticated sessions with the user name as a header, for instance. Our desired end-state - is when we do not have any knowledge of or access to credentials used by the end-user.
To enable us to make this transition possible - for the time being we store the passwords, since as i explained in another reply, we must make this transition transparent to the users and we have legitimate users that login once a year (think credit report, just an example, not our business), so we can't intercept the credentials within a reasonable time period.
/quote
quote:
There's no argument from me regarding undesirability of keeping the passwords. But we also have a different concern - we don't want to be in the business of authentication at all - our goal is to have a third party service or appliance performing authentication and forwarding us already authenticated sessions with the user name as a header, for instance. Our desired end-state - is when we do not have any knowledge of or access to credentials used by the end-user.
To enable us to make this transition possible - for the time being we store the passwords, since as i explained in another reply, we must make this transition transparent to the users and we have legitimate users that login once a year (think credit report, just an example, not our business), so we can't intercept the credentials within a reasonable time period.
/quote
The approach i used applies salt. It can also pad the password to some standard length, if desired. Though it currently doesn't do that.
Slowness of the algo is perhaps a benefit in a way? In my testing it is still fast enough for our use (we are not facebook).
How about this: User provides their public key when they signup. To authenticate, server produces a nonce and sends to the client, client signs the nonce with private key and sends the result, server verifies signature with public key.
What are your reasons to hold on to the password?
I would worry about trying to dual-purpose the encryption like this. If you feel like you must have a way to get the plaintext password back, which BTW you really really do NOT want that liability, but anyway, the safer choice would be to use a semantically secure encryption of the password, and then ALSO hash the password using the current best practices. Use the hash for now to authenticate users, keep the ciphertext if you must, but you'll probably regret it.
There's no argument from me regarding undesirability of keeping the passwords. But we also have a different concern - we don't want to be in the business of authentication at all - our goal is to have a third party service or appliance performing authentication and forwarding us already authenticated sessions with the user name as a header, for instance. Our desired end-state - is when we do not have any knowledge of or access to credentials used by the end-user.
To enable us to make this transition possible - for the time being we store the passwords, since as i explained in another reply, we must make this transition transparent to the users and we have legitimate users that login once a year (think credit report, just an example, not our business), so we can't intercept the credentials within a reasonable time period.
What i try to get out of the replies - is whether there's nothing immediately broken about using asymmetric encryption that endangers us aside from the understandable but not immediate concern about liability.
P.S: Earlier i mentioned "intercepting credentials" in our code - and wanted to highlight that the ease of doing that is precisely the reason why we really want to separate ourselves from handling credentials in any form.
> ...
> The above construction may invite theoretical concerns about entropy reduction (i.e. 72 characters of raw binary without any NUL bytes comes out to about 573 bits of possible entropy, but a SHA-384 hash outputs are clearly limited to 384 bits).
Given BCrypt hashes are a mere 184 bits, I don't see how this is a meaningful concern even in principle. If you're brute-forcing search spaces this big you're no longer looking to recover a password, but find a collision.
This was added in response to a point that a couple people (or perhaps a convincing sockpuppeteer) raised and tried to use to decry the entire article. You're lucky to get 60 bits of information entropy in any given user's password, as is. The "theoretical weakening" here isn't a practical concern: "2^192 security" is still boring crypto.
https://blog.agilebits.com/2011/05/05/defending-against-crac...
PBKDF2 is an improvement over PBKDF1 (and other naive iterated hash constructions), but attacks got better and better defenses are called for.
No, it really isn't.
In reality, PBKDF2 is fine. It's not the best thing you can use, in the same sense as AES is probably not the absolute best round-for-round, cycle-for-cycle cipher you can use, but it gets the job done.
The answer for "what password hash should I use" can accurately be summed as "put bcrypt, scrypt, Argon2, and PBKDF2 on a dartboard, and then throw a dart".
In 1PW's case, PBKDF2 is even more reasonable, because 1PW actually needed a KDF, and bcrypt is not an especially good KDF.
You can create a worker system, or use a child process to solve this problem, but most of these articles never mention it
[1] https://github.com/ncb000gt/node.bcrypt.js#async-recommended
From a quick glance at the source, the answer to the first is no.
I wonder whether threading really saves you, given that there are only so many cores in a system and password hashing is CPU heavy. Naively it seems a DoS would involve sending as many parallel requests as there are cores, which is not a lot and can easily be done from a single machine.
It's also hosted on a relatively cheap VPS.
For things like blog posts, I'm sad that Movable Type fell out of favor. The pattern of using dynamic server pages to edit blog posts and generate static files for your viewers just makes so much sense to me. Then again, these days we can host our single-page JS front-ends on S3 + CloudFront and hook them up to API gateway and Lambda / Beanstalk talking to RDS and "not have any servers" :)
I dont think it could get a real issue in a real environment, except someone really wants to fuck with you. But the some fail2ban on to fast requests should fix this easily as the hashes dont actually take that long.
Using `php -S 0.0.0.0:80 index.php` would do it!
Not serializing IO, but them stopping every worker just because one of them has some hard work to do is not a sane working model for web backends.
But yes, I agree with you that that is why every web app _should_ be threaded.
The one recommended in the article[0] actually includes a blowfish C++ binding, with both synchronous and asynchronous methods for comparing and hashing. That binding uses NAN[1] to queue AsyncWorkers. NAN uses libuv[2] behind the scenes to launch a thread[3].
0. https://www.npmjs.com/package/bcrypt
1. https://github.com/nodejs/nan
2. https://github.com/libuv/libuv
3. https://blog.scottfrees.com/building-an-asynchronous-c-addon...
Unfortunately, the example used for the bcrypt node module is the async form of hashing. That library exports both sync and async forms of hash/compare.
Anybody here that knows if any of these is a valid replacement? https://nodejs.org/api/crypto.html
Or you can do the bcrypt() inside your (PostgreSQL) database and add an extra layer of security by not directly exposing the hash (or even the whole password table) to the web server process (i.e. write a password check SQL function and allow only calling it, no direct table/column access).
I did exactly what you are suggesting on a project some time ago - deferred to the database to do the password hashing and comparison to the stored hash. Unfortunately, the server query logs contained the query parameters. So did some of the database logs when running at elevated log levels. We quickly decided that was unacceptable, and moved the hashing process into the web server.
The webserver framework I was using at the time logged all its queries to the database at the normal log level. It couldn't log an actual query string because we used prepared statements, but it reconstructed the query and substituted the parameters in.
Enormously helpful for debugging an issue, when you can just copy and paste a query out of the serverlogs. Although I would probably have preferred that it not do that at the normal log level.
Curiously, it never logged the http request string/parameters. When I wanted that I had to add my own code to the request handler.
These kinds of systems aren't technically very complex but there hasn't been a lot of traction for them outside of corporate environments. The result is that they tend to be hyper-expensive, unfortunately. Standards like FIDO and the Yubikey are good first steps at pushing this into the consumer space, although they don't offer on-device PIN validation yet.
The worst part for me would be loosing that thing, in the end i would need alternative login methods anyway to be sure i dont lock myself out.
Classic Authy on a Smartwatch would be the simplest method i could live with that comes to my mind.
Alternative Login is always a problem, their is always a tradeoff between security and useability. You could easly print out some backup access codes. You can continue to use your E-Mail as a anker, you get an E-Mail and then your allowed to register a new token. That leaves the question of how does your E-Mail provider secure its login? I think Google is a good example of the options that are possible.
I don't have a solution myself, but I hope there is one some day.
New Samsung phones and Lenovo Laptops already support this on the fingerprint sencor but the true benefit is that you can have competition among local authenticators without the server having to change anything.
For legacy devices where you do not have any kind of local authenticator their could be local software that lets you enter passwords. Its already part of the standard that the server has a way to only allow some of the authenticators, based on certification. This would allow a bank to specifiy only devices manufactured by a specific company or that have passed some certification process.
Now some people of course say that fingerprint sensors are not very secure. That is true. It still increases security because the local authenticator will only sign the challenge if it is sent from the correct website/software (App Id) and in most cases the TLS Channel Id. UAF (and/or U2F) prevent far more common attacks at the cost of less security if an attacker actually steels your phone. You can combine UAF with U2F for additional security even if your phone/laptop is stolen.
This has been passed to the W3C and its hopped FIDO 2.0 will be an offical standard [2] [3].
Google, Github and Dropbox already support U2F. PayPal supports UAF (works on mobile). Both of these can also work with NFC and Bluetooth LE.
[1] https://fidoalliance.org/specifications/overview/
[2] https://fidoalliance.org/fido-alliance-announces-fido-authen...
[3] https://www.w3.org/Submission/2015/SUBM-fido-web-api-2015112...
Is there something obvious I'm missing here?
hmac.compare_digest is a constant time compare, in that no matter if there is a match or not, it will take the same amount of time.
F("value") > "123455" which is close, but that does not let you get a 'better' guess.
PS: Assuming the Salt is hidden, and the Hash is secure.
So if the timing attack has told you that the first letter of the hashed version is 'A' then you find another password from the table that hashes to AAxxx, ABxxx, etc.
Obviously that depends on you being able to precompute all the passwords in a consistent way as the hash being used.
Along the same lines - could you use a timing attack to figure out a salt? I guess it's near impossible?
I was going to develop this into an exploit tool, called TARDIS (backronym for Timing Attack to Remotely Dispel the Illusion of Security) against, e.g. Piwik, Oxwall, and other products that still use MD5 passwords. The main reason I didn't was: No free time to build it and tune it against the internals of various programming languages' == implementations.
Using == instead of hmac.compare_digest is unlikely to be a source of vulnerabilities in your application, but it's a good habit to get into whenever you touch cryptography.
See also: https://news.ycombinator.com/item?id=10345965 (pg. 33-36, 42-43 of the PDF)
Again consider timing leaks.
Many interesting questions:
How do secrets affect timings?
How can attacker see timings?
How can attacker choose inputs to influence how secrets affect timings?
Et cetera.
The boring-crypto alternative: crypto software is built from instructions
that have no data flow from inputs to timings.
Obviously constant time.If the attacker knows which algorithm and work factor you're utilising and your system doesn't use randomly generated per user salts (or an unknown pepper) then theoretically an attacker could use a hash timing attack, combined with a rainbow table, to massively reduce the scope of a user's potential password.
For example, let's say your rainbow table has 20 million password-hash combinations and your hash length is 33 letters, for every letter I know I could drop literally millions of hashes I know it ISN'T. With just the first four letters I could drop it by almost 60% (although my maths here might be completely wrong, it is actually pretty complicated to determine). But there's no real limit on how many letters you could drop from the hash with a timing attack, you could turn someone's password into a 1 in 62 chance.
The TL;DR: Use a per user salt. But a timing immune comparison definitely is "defense in depth" in case there is a bug elsewhere that breaks salts.
Password hashing has existed since at least the 1970s and until the 2000s salting wasn't common. While some hashing libraries do insist on you supplying a salt, it is still ultimately up to the application developer to generate and store the salt for later usage.
Therefore it is still common for an application developer to "worry about salts" even if just for storage and generation reasons.
Anyone using 3DES, MD5, or similar is likely vulnerable to timing attacks, and they're definitely still in the realm of a "password hash." Plus some wonderful developers hard code the salt (salt = "secret") which too could leave them vulnerable to timing attacks if an attacker knew what hashing algorithm and workfactor (e.g. the default) was in usage.
I truly hope this will help then: https://paragonie.com/blog/2015/08/you-wouldnt-base64-a-pass...
"Password hashing" is its own compound noun. The acceptable algorithms (as defined in the blog post this HN thread is about) take care of this for you.
> Anyone using 3DES, MD5, or similar is likely vulnerable to timing attacks, and they're definitely still in the realm of a "password hash."
3DES is a block cipher. MD5 is a crytographic hash. Neither of them are password hashes.
> Plus some wonderful developers hard code the salt (salt = "secret") which too could leave them vulnerable to timing attacks if an attacker knew what hashing algorithm and workfactor (e.g. the default) was in usage.
What you're describing is closer to a "pepper".
http://blog.ircmaxell.com/2012/04/properly-salting-passwords...
3DES is both a block cipher and a cryptographic hash. At least UNIX thought so in the 1990s as many MANY people were storing UNIX passwords in 3DES, DES, MD5, and similar.
> Neither of them are password hashes.
25 years of computing history would disagree with you. MD5 was the defacto standard for password hashing for almost fifteen years.
But no doubt you'd playing silly word games, and are going with your own definition of "password hash" that includes or excludes different hashing algorithms as it is convenient for you. I won't get drawn into that.
> What you're describing is closer to a "pepper".
What you're doing is called being "condescending." You know full well from my posts above that I am familiar with salt/peppering/hashing, and the different technologies involved. So linking to 101 tutorials and definitions of basic terms is only intended to aggravate.
If your password storage mechanism requires you, the programmer, to generate a salt, you may well be using the wrong password storage mechanism, or using it in the wrong way.
… then you’ve already lost, constant-time comparison or not.
Also, rainbow tables are obsolete.
I know this doesn't apply to banking, etc but 99% of the websites that "require" me to create an account and log in don't need to store primary credentials for me. Please pick a secure implementation of oAuth2 and let people store their credentials wherever the hell they want to.
I'm bored of getting hits from "Have I been pwned?"
> 99% of the websites that "require" me to create an account and log in don't need to store primary credentials for me
Why are you giving them valuable credentials? Give them a throw-away password (password managers are great for this).
That way users can be their own oAuth providers if they want.
It's hard to make a blanked recommendation like that, even for "only 99%" of websites. Neither you, nor the person building the website, has any insight into who the website's users trust.
Offer OAuth2 as an alternative to passwords: Great move.
Only offer OAuth2 and don't let people create an account: Questionable.
They can host their own.
I don't understand why they would trust <crappy forum owner> over a dedicated authentication storage place but that's their choice. And yes, there is also every possibility to offer direct credentials, per the Stack Exchange model (they host their own oAuth server and allow simple registrations).
What if <crappy forum owner> happens to be a security engineer, and <crappy forum> happens to be Silk Road 13?
The trust decisions people make are situational and nuanced. OAuth is great if that's where people invest their trust. Otherwise, you're outsourcing it for the user to a company they might fear.
1. Let every website on the Internet potentially be an OAuth provider.
2. Make OAuth optional.
If you follow option #2, then this article is still relevant because you need to handle passwords securely.
Secondly, every website on the Internet is potentially an OAuth provider.
Not to mention that I have —on multiple occasions here— suggested that websites that consume OAuth should also provide it (like Stack Exchange).
A hybrid between the two (common OAuth-style endpoints and any OpenID endpoint) is the best solution for everybody.
The "bindings for most programming languages" link goes to the libsodium documentation, but I'll add a link in more contexts.
There would appear to be three distinct Java bindings of libsodium.
https://docs.djangoproject.com/en/1.9/topics/auth/passwords/
Passwords should be scrypt'ed on client, and then, the server should generate a SHA256 hash of the scrypt'ed hash and store that in DB.
- Running CPU & memory heavy scrypt hashing on the client side will allow us to use bigger hashing work-loads.
- EDIT: Removing the MITM point, because as many said, that's the job of TLS anyway.
- External brute force attackers will have to take the burden of heavy hashing. No DOSing through scrypt.
- Storing SHA256 hash instead of scrypt hash on DB means even if DB is stolen, attackers can't use stolen scrypt hashes to authenticate any client.
I would love to get others' feedback on this. EDIT: Found the reference: https://news.ycombinator.com/item?id=9305504
Doesn't that scrypt hash then become the password, from the server-side application's perspective?
How are you storing the salt for the user if your server only knows about a SHA-256 hash.
> MITM attacks won't get access to unencrypted fields.
That's TLS's job. If you, for example, are building a web app and you're delivering Javascript to perform the scrypt calculation, a MitM can replace the code to exfiltrate the user's plaintext password. It doesn't make sense for the threat model.
Yes, the scrypt hash just becomes the password for the server. However, this password is going to be unique and almost impossible to guess. Salting this password on server side doesn't give us any additional benefits. The server can store the salt that the client used, though.
Yes, the MITM point was minor. Just another minor security benefit on top of TLS (which takes the majority of the burden of securing against it). So, maybe I shouldn't even mention this point.
I don't even your threat model.
Can you walk us through how you see this working?
Client takes 'crappypassword' as input, sends p = scrypt('crappypassword')
The server stores SHA256(p)?
EDIT: misread what you said. is this correct?
Otherwise, you're relying on the user side to properly safeguard bcrypted password. That is often wrong. You cannot store any kind of transformed password anywhere, or it's no longer a password, but a token.
Yes, those "Remember me" things transform passwords into tokens. Tokens can be stolen.
I'm going to refer to bcrypt because that's what your comment used, but the parent post used scrypt.
> Are you assuming that collisions against SHA256(bcrypt(p)) are harder than SHA256(p), right?
I don't think it is. I think it's assuming 2 things:
1. That SHA256 collisions are rare enough, and SHA256 attacks are hard enough, that they can be ignored.
2. That the output space of bcrypt is large enough to overcome the risk that salting is usually designed to overcome.
On the first point: This approach simply isn't worried about SHA256 collisions between passwords. Yes, it's theoretically possible that 2 users with different passwords and/or different salts might end up with the same hash. If that happened often then it would be an issue - if you had n distinct users in your system but only n/2 distinct hash values, then if an attacker had a copy of your password store, it would effectively double the pay-out each time they successfully cracked a user's password (that is, each cracked password would allow them to authenticate as 2 users).
But in practical terms, collisions are going to be tiny, and you're really worrying about the case where n distinct users have n-1 or n-2 distinct hashes. That's not going to meaningfully change your exposure.
What is more of a risk (and this might be the point you were making) is that if any of your hashes happen to collide with the hash of a known input, then you're screwed, and while that's unlikely, "unlikely" isn't an ideal protection.
When you're storing { SALT , HASH( SALT || INPUT) } you're reasonably protected against such collisions because they would only be effective if the known input happened to start with the salt.
What the bcrypt approach can offer is that it knows that the input to SHA256 needs to be the output of bcrypt, so you can refuse to accept anything that isn't 186 bits long (or 31 base64 characters, depending on your approach). That constraint might well be stronger protection than a salt, although I haven't run any numbers.
Probably adding a salt is safer, and I was going to attempt serious analysis on this scheme I'd certainly want to test whether that was true or not.
On the second point, salting is usually designed to work around the scenario where 2 users have the same password and therefore (absent a salt) would have the same hash. It is not specifically intended for the scenario (described above) where two users have different passwords that happen to hash to the same thing.
In the approach discussed here, the input to the hash is the output from bcrypt. That value has already had a salt applied, so 2 users with the same original passwords would be providing different inputs to our hash function, so we would be storing different outputs.
> you're relying on the user side to properly safeguard bcrypted password
I think that's a genuine issue. You need to do a not-insignificant level of client-side processing on the user's password. You're relying on the browser disposing of that data securely. It's "just bits in RAM", but bits in RAM leak, and there's no way for the browser to know that the bits you were working with were sensitive.
Not sure how I feel about it.
How are the salts managed? This detail is important. SHA256(scrypt(password, constant_value_instead_of_salt)) is going to produce collisions in the stored hash.
I think it's dangerous because of the dependence on the client, and useless because scrypt(password,salt,work_factor) on the server is plenty hard to attack.
It enables DoS attacks, though, right?
For a large enough work_factor. In practice "large enough" is usually interpreted to mean "something that seems safe without being so large that I need to buy lots more servers"
The argument (which I'm interested in, but not yet sold on) is that moving the scrypt to the client allows you to pump up the work_factor even higher than you would have been willing to do on the server.
In general though, pushing auth down to clients in Javascript makes my skin crawl. You're one XSS away from having an attacker no-op your scrypt and return SHA256("secret-attacker-password") on registration. The logical follow-up to that: 'have the server run scrypt the first time' -- but then you've just moved the tough work to user registration, which seems just as exploitable.
I dunno. Safe crypto is hard enough already; I'm not sure pushing it into Javascript on the browser makes it any easier.
The DOS risk on registration can be mitigated more easily than the login one. e.g. Rate limiting new registrations will often be more palatable than limiting logins, or you can require "email validation" before you set the first password. And, not every application allows self registration.
On mobile devices client side hashing could affect battery life. If done with javascript, then login won't work without javascript enabled. Also, a work factor that doesn't slow down the slowest devices that a user might use could provide much less protection than desired, although the server could also do some work.
Before you mention the salt - the salt is only useful if it if available. The client needs to know the salt in order to perform the calculation, therefore your attacker will also know the salt, so that doesn't count as entropy in the hash, either.
base64_encode(hash('sha384', $password, true))
In client side JavaScript. I've seen my own passwords scroll in front of my eyes when debugging servers and reading POST variables. It's a minor level of shoulder surf protection before it hits the proper hash on the server.Precisely the reason why password input fields usually display replacement characters instead of actual characters (or nothing at all if you're on a Unix terminal).
A MITM would let you hijack the JS that controls scrypt/SHA-256, so you're already at game over. You've got to deliver that to the user in some fashion (TLS!); this isn't a real win for your approach.
> External brute force attackers will have to take the burden of heavy hashing.
If the attacker is trying targeted access to a site the only thing that's relevant is the time they have to expend - the hash is opaque. If they have the hash from, say, a DB theft, they're already going to have to take that burden. Your approach doesn't seem to add anything.
> Storing SHA256 hash instead of scrypt hash on DB means even if DB is stolen, attackers can't use stolen scrypt hashes to authenticate any client.
I'm not sure what that means. If you steal my scrypt hash for example.net, how do you use that to authenticate me to example.net?
Your approach pushes out complexity to the clients and doesn't seem to win us anything.
algorith | fairly safe difficulty (all variables) | very safe difficulty without incurring too much performance cost.
pbkdf + sha1 | completely unsafe | completely unsafe
pbkdf + sha2 | 100000 | ...
pbkdf + sha256
bcrypt
scrypt
argon2
That's the problem with a chart like this. The gradation will go from "completely unsafe" salted hashes to "very much safe enough" with only marginal changes after that.
Another problem is that these functions are all parameterized, so the chart needs to capture the safety level at specific parameters.
algorith | a safe set of parameters | a VERY safe set of parameters but needs strong hardware
Maybe a 4th column of "unsafe if difficulty is below"
This could be something updated yearly or whatever to help people figure out what things they need to move towards. If they see that their currently used settings are below the unsafe line, they will know it's time to upgrade.
Example: I use PBKDF2-SHA1 at difficulty of 11k iterations. Is that number still within the "probably okay" list, or is that in the "I can hack any password on your list in 15 minutes"
[1] https://blog.afoolishmanifesto.com/posts/do-passwords-right/
1. Use the best option available, but
2. We provided example code in multiple languages for the
best one that's widely availableThere is an open issue in Coda Hale's bcrypt repo about this: https://github.com/codahale/bcrypt-ruby/pull/119
My stance on posting "best practice" articles is: you must follow all best practices in them.
<Edited for clarity>
Unfortunately, there's nothing I can do about that, unless someone can point me to an alternative that uses a constant-time comparison.
I've left a comment on the pull request so that, hopefully, it can be merged.
> (and likely others)
Which others?
I see your team is actively posting on the bcrypt-ruby issue #119 as we speak, so I guess I'd say wait for the PR to merge, or manually implement the approach they've outlined to secure compare: https://github.com/codahale/bcrypt-ruby/pull/119/files
Timing attacks depend on an attacker having control over the hash being compared (e.g. they have a HMAC in a cookie they're sending you, and they can adjust it character by character) - with randomised secret salts and server-side hashing, this isn't the case.
The SCrypt example is the one I'm more concerned with. The defaults there are 1MB of RAM and 64-bit salts, both of which could do with increasing. I have an open issue on this: https://github.com/pbhogan/scrypt/issues/25
As it is, the example would probably be better as this:
password = SCrypt::Password.create(usersPassword, salt_size: 32, max_mem: 16*1024*1024)
You'll be pleased to know SCrypt::Password#== is at least constant-time.It's worth spending the time to learn, implement, and integrate into the security best practices you should already be deploying.
Of course the core of it could be upgraded, but the idea is sound. Sounds like you could say the same thing about using a DES-based hash. "Don't Hash! It's weak!" is throwing the baby out with the bath water.
Sure, public-private pairs are more difficult (to impossible) to brute force once they're so long, but srp, bcrypt, pbkdf2 &c are pretty good when you want/need a password. Srp being superior in that your secret never leaves your computer.
SRP gets you (1), but is inferior to every other modern password hash on (2). SRP is better on (2) than other PAKEs (it's an "augmented" PAKE because it tries to slow down brute force), but password hashes have (2) as their whole objective, and total design freedom.
You can combine PAKEs and password hash concepts. But for the purposes of password storage, SRP isn't buying you anything (except for a bunch of possible crypto bugs that will gameover your project).
SRP is a bad call for virtually all projects. Or, maybe a better way to say that is, "if you have to ask, don't use SRP."
Only until that was what people cared about. How long, and how many still, use single-round md5 and sha for password storage?
SRP is an old algorithm -- it could be strengthened significantly with little love. (Similar as to how took a little care from the community for bcrypt, pbkdf2, &c have become the "norm" over md5.)
> SRP isn't buying you anything
SRP is buying you the ability for the server to never have the password, a secret value you're handing someone else!
> SRP is a bad call for virtually all projects. Or, maybe a better way to say that is, "if you have to ask, don't use SRP."
I still feel like you're throwing the baby out with the bathwater. Just because something can be updated and hasn't (like using a weak hash like md5), doesn't mean the core concept is bad (using a 1-way hash function).
Most of this can be solved with SAML, but it requires all of these old enterprisey stacks to convert their auth systems to be SAML compatible.
In the long term though, I think SAML and OAuth will be more prevalent than using passwords everywhere.
Are there any OAuth providers established that simply provide email addresses? I feel there's definitely room for a company to setup an OAuth provider that doesn't do anything except provide identity, rather than relying on personal information harvester companies (FB, Google, etc). The hard part would be getting onto the approved OAuth providers list of all the sites that users want access to.
There's also the fact said systems seem to be a nice target for spammers. They're popular, so they're often attacked. And because they're often attacked, their anti spam defences don't usually last very long. So any spammer now has a nice way to get an account on near enough any site they like, while with a standalone system, they'd at least have to tailor their attacks to the site in question.
There's also the fact it silos much of the internet (or at least user information) within the systems of a few large providers, which gives a significant amount of control to said providers.
Yep, I do tailor the selection of oauth providers to my audience. E.g., an app for devs = google, github, and gitlab.
Don't. Unless it's 100% absolutely necessary.
If you must, continue reading on.
People who know about password hashing don't need this advice.
People who are storing passwords need this advice.
For many services, the most valuable data they contain happens to be the user credentials that the service uses to authenticate the identity of the user. If you assume that users share their credentials across multiple services, if your system is attacked, and you improperly stored their credentials, you've caused way more damage, than the data you were trying to secure.
For example, take your run of the mill online todo app. If an attacker got access to the all the todos, vs an attacker got access to the all the todos and all the user passwords (even if securely stored). What is more valuable to an attacker?
Generally: Being able to deploy malware through a site users trust so you can attack them directly, en masse.
But the password hashes are a good runner-up.
For mobile apps, I think using SMS and One-Time Passwords is sufficient as well.
You can also use a delegated auth system put in place by OAuth v2, or something like Google Accounts, Twitter, etc. Those services do get pushback, as not every uses Gmail (or likes to give any run of the mill app access to Google account data).
What do you mean by privileged elevation (I'm talking in the context of a web app)?
> just using an email address and send "magic links" + persistent (but expiring) sessions
Do you mean something similar to what Medium does?
https://medium.com/the-story/signing-in-to-medium-by-email-a...
I like it too. But what about delays in receiving emails (when I suggested this approach to some customers, they were worried about that)?
> You can also use a delegated auth system
Yes, outsourcing the authentication is another solution, but some (most?) users are not comfortable with giving so much power to Google/Twitter/GitHub/etc. (as you wrote in your comment).
For example, prior to changing account settings, reauthenticate the users, regardless if they have a valid session.
> I like it too. But what about delays in receiving emails (when I suggested this approach to some customers, they were worried about that)?
Combo of email & SMS.
As Bruce Schneier has written, security is usually a trade off of user experience.
Why not reauthenticate by sending a one time password by email/SMS (instead of asking for a password)?
> Combo of email & SMS.
Good idea. I was also thinking of sending one time passwords through some chat applications like WhatsApp, but most of them have no API for such a thing (except Telegram).
In France, where I have most of my customers, the SMS is paid by the sender, which makes it expensive.
(Title is currently: "How to Safely Store a Password in 2016")
People used to memorize whole books.
is that true?
Proprietary Analog Password Encryption Routines.