All you ever wanted to know about building a secure password reset (NSFW)
troyhunt.com
troyhunt.com
If someone attempts a password reset on foo@example.com, and there is no record in the DB for foo@example.com, immediately tell them so.
Why?
1) While this does "leak" the fact of foo@example.com not being in the DB (and, by consequence, can be used to verify whether any email is in the DB), naive implementations of the best practice will also leak that via a side channel. If you would not immediately greenlight $20k to have a security consultant spend one week auditing this feature extensively to avoid the timing attack, a) your app is not important enough for this level of security and b) this buys you no real security, just the illusion of it.
1.5) This is totally security theatre if you have a registration process which checks for provided emails already being the in the DB, for example for telling people "That email is taken." Different route through the maze but same cheese at the end.
2) You can already spearphish people for their accounts on example.org . Start with an email list of people. Spearphish for their accounts irrespective of whether they have them or not. Collect wins since losses cost your botnet nothing at the margin.
3) MOST IMPORTANTLY: you're buying yourself six years of frustrated emails from customers which had 90% totally automatically discoverable solutions if you do this the "right" way. I'm speaking from experience. Many of your customers have multiple emails they habitually use. If you do the best practice, every time they misremember which one they signed up with, you get a support email. This will be your #1 cause of support emails and the cost for dealing with them is astronomically higher than the minimal marginal gain in security for your users.
Edit: Omnibus response for people doubting that a side channel attack is possible:
The attacker wins with microsecond precision over the open Internet and nanosecond precision over the local network, which is plausibly achievable on all of the cloud providers than HNers like to host their apps at. Microsecond precision is enough to discriminate between "record in DB" and "record not in DB" on many plausible application architectures, even if your HTML/headers returned are exactly identical for the two cases (and they frequently won't be out of the box).
For example, the relevant code for Bingo Card Creator includes the line
user = User.find_by_username(params[:email])
This line takes approximately 30x longer to execute if it finds a user in the DB versus not finding a user. Why? ActiveRecord. You could, if you had a mind to, trivially discriminate between the two cases in under a week over the open Internet. If you really wanted to ruin my day, you could get a prepaid Visa card and get a VPS from Rackspace on the same intranet, then discriminate between the two responses in under an hour, again without tripping anything.
Some folks think there are trivial ways to defeat timing attacks. You are either mistaken or you have a very different view of the word "trivial" than I do.
However, the original article did raise a point - there are some contexts in which leaking membership might be especially embarrassing. I don't think any of my coworkers would use timing attacks to determine if my email address is in the database of someembarrassingpornsite.com, but typing in my email address and seeing if you get an error is within the intuitive grasp of many preschoolers. Perhaps the guidance should be that for 90% of sites you don't need to worry about disclosing membership, for 9% you should consider a naive implementation of the best practice, and 1% should get in a security consultant.
Have you considered doing what EDW did and compiling a book of your HN replies, expanded by further insights gained while doing so? Would you have time? I'm sure many would find it valuable.
int MAX_RESPONSE_TIME; start = Date.Now(); user = Database.findUser(email); SleepUntil(start + MAX_RESPONSE_TIME); return user;
Note: just doing this in your app isn't enough, because you run in a framework and on a virtual server, and things like memory usage affect the speed of these components.
This is specifically a security problem. Saying "I'll sweep it under the rug" and considering it solved is historically a bad idea.
This chain of "I'll just do X", "no, X isn't enough", "then I'll do Y", "no, Y isn't enough", "then I'll do Z" are typical of these problems.
Unfortunately, each "I'll just" is code for "I'm not going to think through this problem because I think it's easy." As with X and Y above, Z is likely to have serious problems.
Example problem with the queue approach: hacker puts in lots of password reset requests for his own domain(s) and sees that the queue workers have delays of 0.22 seconds between sending back responses to his domain, with gaps which add 0.37 seconds or 0.22 seconds, to within some error. Inference: a wrong email takes slightly longer because it often has a retry, and the 0.22 second entries are other correct password reset requests to valid domains.
Now he can send back a stream of passwd requests to his domain interspersed with three or four requests to a given email address. Now he knows if it's valid or not. ("I'll just block his domain?" "He'll use a botnet, he is anyway" "I'll just rate-limit password resets!" "Good idea - now, do you really believe that that was the last hole in your swiss-cheese security?")
Now you could fix this, too. But see how we've gone through yet another "I'll just?" You can make sure to do the delay in lookup and email sending on each entry in the queue (by the way, you've slowed down all your queue workers and likely increased your hardware costs), but trying to not think about this and sweep it under the rug is really hard.
When patio11 says that timing attacks aren't trivial to fix, this is what he means.
Is the "put it in a queue, add a delay to the queue" method perfect? That is, is "I'll just do Z" the answer? No. Having the queue sleep() is less load on other components, which an attacker can check for by pinging other services. Having it busy-wait is a different load profile than receiving an email. That's the microsecond-precision thing that patio11 mentions where an attacker can get another instance on the same Rackspace network and measure with a side-channel attack.
This is a really hard problem because there are an almost unlimited number of side channels here, and your OS doesn't try to keep information from leaking out on them.
That would be relatively expensive, but if you just do it for login/registration/password resets it might go a fair way towards mitigating those sorts of attacks.
But yes. Every layer you add (maybe) makes it a bit harder for an attacker. You simply can't stop a dedicated attacker, reasonably speaking.
"Do not respond before" would at least make an attacker have to use a (somewhat) less reliable second channel to find out. Expensive, but it would do you a bit of good against casual attackers.
Beyond that, you can't really block unless you block all the channels -- i.e. add a "do not respond before" to everything, not even just web requests.
Also, when you say "do not respond before" is a header, I assume you mean set by your reverse proxy before the back en handles it. Clearly setting that from the client won't do you any good at all :-)
The concept was just that no request for paths related to logins or passwords would take less than an amount of time, eg. 0.1s or 0.5s. It could even just be a config option.
Configuring it at the firewall/web server side would be an easy way to make life harder for an attacker, without having to fiddle with (or even understand) the internals of a web app.
That's why this is hard. You don't get to control what channels of information the attacker looks at.
And while you may think that the fact that your app isn't a risk because it doesn't deal with sensitive issues, your users may not agree.
Sometimes security is important. Other times its just a waste of energy.
Site's registration form validation will give the same information and without side-effects.
Can you explain this to me? Is the only difference the error message used, or are you suggesting that apps should support multiple accounts with the same email/username but with different passwords? Or something else entirely? :)
Personally, I don't understand what difference a message makes; it'd be trivial for a user to test by signing up for an account themselves. And not having a unique login identifier would be a nightmare. Which makes me think I'm missing something obvious (damn this headcold, it makes me feel so much more stupid than usual!).
If you were to try to register a new account with an email address already in the database, you should get some variation of "A user account with that email address already exists," verifying the existence of an account with that email.
But the same trick works for signup forms as well: I could get an email to my mailbox "hey, thanks for registering again but you already have an account. If you forgot the password, click here." That plugs that hole. Note that this won't work with usernames, but usernames are far less tied to a specific person than email addresses. (apart from some people that have very specific and well known usernames).
But how would that work for users who aren't you? My name is - surprisingly, to me - quite common, and the number of registration signups I receive at my [firstname][lastname]@gmail.com email address is really quite surprising.
So if you suddenly get a flood of "hey, thanks for registering again" mails you'd at least know that somebody is trying to tamper with your account. The email could even say so and add a "please notify us if you think somebody is trying to play tricks on you" link.
Nobody else could have registered with that address and (rightfully) expect a confirmation email.
I regularly receive confirmation emails from websites where the user believes their email address to be john.doe at gmail, instead of johnathan.doe at gmail. If this is common enough for my name, it must be really common for more popular names.
So, following your example through, john.doe receives the "Hey, you're already registered!" email, and johnathan.doe first thinks they have registered successfully, and later on thinks that my service sucks because they can't log in, reset their password... and registering appears to do nothing at all. User confusion - and support headaches - ensue.
So the variant of always sending an email and always accepting the registration provides the required benefit with a minor drawback.
[1] Unless you don't send confirmation addresses at all which would be pretty much illegal for most services in germany since double opt in is required for pretty much everything of interest.
edit: Since this was regarded as a statement on legal matters I herein clarify to mean "pretty much anything of interest": I loosely intended to say "most things a commercial service might want to do with data, including but not limited to sending me emails which might be regarded as an offer or an incentive to buy any paid service or any promotional email." As has been stated further down it's not a legal requirement to confirm email-addresses in all cases.
Ah OK thanks, I understand now. (have a headcold that is confusing me right now, so if in doubt, it's my fault ;)
I think that the only thing were quibbling about is what a "minor drawback" is to each one of us. For me, it's not such a minor issue, but it's been an enlightening conversation with you, so thanks :)
I agree. But that's always the case with security and I think in this case you can easily fix the drawback with a clear messaging such as "This is what you entered: (replay form data). You should receive a confirmation email within (x) minute. If you don't make sure the email you entered is correct." You'll need that message anyways to catch those users that enter a completely false email address anyways.
[1] Most services that I've signed up to lately log you in once you confirm the email. That's what I regard as the best compromise and for that case, the scheme works perfectly without leaking information.
The second is the usability question. Sites where you can't login until you have received and clicked the activation link are throwing away signups. The usability of "wait a half hour before you can do anything" is really, really poor. You can certainly argue that having fewer signups is a worthwhile trade-off to gain some privacy, and in some cases I might even agree. But I don't think that is true in most cases. As others have pointed out, you get thousands of emails about signup/login/reset related issues when you try not to leak this info. You get zero emails about leaking it.
There's other reasons to use double opt in. I register for your service with no double opt in and I have a typo in my email address. I then log out and forget about it. I just lost my account. Double opt-in prevents that. Think of a forum where you can register with an email address and make public statements - if said forum has no double opt in and you register with my address and slander someone I'd take that forum to court since they neglected to prevent that. I might not win, but the forum would be drawn in the fight.
I know that most corporate lawyers I've worked with get twitchy if you propose removing double opt in - even in cases where it's technically not required. I guess lawyers are more the "play it safe" kind of people.
I agree with you that double opt in is not the silver bullet that magically fixes everything, but as I said - we're deep in trade-off territory here.
I also dispute the point that you get zero emails about leaking the information that someone is registered. I have worked on projects where that information was absolutely privileged and it was of utmost importance that no info about who's registered could be leaked.
Your view of the legal situation is laughable. You are welcome to do whatever you like, but don't try to claim it is a legal requirement unless you are going to back that up with facts.
>I also dispute the point that you get zero emails about leaking the information that someone is registered. I have worked on project where that information was absolutely privileged and it was of utmost importance that no info about who's registered could be leaked.
It doesn't seem like you are trying to discuss this in good faith. Read my post again, I was pretty clear that privacy matters in some cases, but that I do not think it is the common case.
I am actually trying to discuss in good faith, but you're last sentence in your post is:
> "You get zero emails about leaking it."
That's what I dispute. If information would have leaked on that project I'd have had a very angry email from my customer in the inbox. Probably rather a written letter in the letterbox ;)
I guess basically we're both in agreement. If you go back an re-read my statements, I do agree with you that double opt in is
a) not the golden end of it all and the one-size-fits-all approach won't work b) and legally not required in some cases (though we differ on how many cases there are)
However, I argue that im my experience most projects will end up with double opt in because
a) they're legally required b) or they might be legally required to do so in the future (like when they plan to send advertisement emails) c) they have risk-averse stakeholders that want every anchor they can have in a (potential, probably imaginary) lawsuit that some bone-headed user might trigger.
In any case you're kinda missing my original point: The starting point of the discussion was not that you're required to have a privacy protecting signup scheme. My only point is that it's possible to have one. If you don't need one, that's fine with me.
Because you are taking it out of the context of the implicit "for a typical web app" that had already been established in the previous post.
>I do agree with you that double opt in is
You are still arguing a false dichotomy of "double opt in" vs "not double opt in". Double opt is entirely irrelevant. The only time I mentioned it was pointing out that it is not in any way a legal requirement.
>My only point is that it's possible to have one
Nobody said it wasn't possible. People said it is a huge usability flaw.
I don't get that point. What's a typical web app? Most of those that I've built had in some way or another email connectivity. Many had a newsletter component somewhere that was used to inform users about new features/offers/whatever promotional content. Even more had the tentative idea of at least keeping the option open. And if you do that, you need to confirm the email address. So probably we differ on the notion of "typical" here and I guess that's a point that can't be resolved.
> You are still arguing a false dichotomy of "double opt in" vs "not double opt in". Double opt is entirely irrelevant. The only time I mentioned it was pointing out that it is not in any way a legal requirement.
Sorry, you kinda lost me here. I don't understand what point you're trying to make.
> People said it is a huge usability flaw.
That's the whole point. IMHO it isn't that "huge" when you have double opt-in anyways. And as I pointed out that there are some reasons to have double opt-in regardless of legal requirements as well, in fact, most services that I signed up for use it. That might be different for you, but it's certainly not a minority or a freak occurrence if you encounter some service that uses double opt-in. So it can't be that bad either.
I fully acknowledge that you have a different view here, that's completely fine with me.
This is what I have been saying for the whole thread. And you are only getting worse. I do not know how I can make myself any clearer, sorry.
Step 1: Determine how long a query for a not found user may take... Let's say 1-2 seconds since there may be many rows to deal with.
Step 2: Multiply the value in step 1 by two or at least add a comfortable margin, perhaps 5 seconds total.
Step 3: On the start of the request performing the look up spawn a separate timing thread.
Step 4: At the end of the request join the timing thread.
This will tie the processing time of the request to the uniform time out value instead of your database processing delay.
Let's say a valid login takes 4.9 seconds, and you are able to rig an invalid login to take 5.0 seconds. You might think that is an imperceptible difference. But a determined hacker with a week can try your script a billion times, and get to statistical certainty that the email doesn't exist.
Not sure WHY they'd ever want to go through so much effort for that info, but it's always going to be possible.
Trying to make two code paths take the same time without waiting on an external timer is deep voodoo and should be left to experts.
Shouldn't the best practice of sending emails include using at least some sort of message queue? You're probably right that most web apps will use a synchronous call for this, which is easy to run a timing attack against, but it should be easy to fix this without spending $20k on a security consultant.
(btw: 30x longer is not of interest for timing attacks. If it's 1ns vs 30ns then it's still 30 times longer but doesn't buy the attacker anything when he can only measure ms precision)
Which is why you are you not really adding any security with this stuff by just plugging the obvious holes.
Also, you have to fix this problem every place else that you touch the user database, such as your signup process. There are normally many places you would be affected by the user database, because often the service you provide is slightly different for each user.
But I'm still not convinced that's easily measurable in practice, considering most sendmail implementations dump messages into a queue rather than trying to deliver them immediately.
[1] "If you're registered here, you might be registered with a different email address. If you're not, ignore this one. etc."
Just to be clear, I think the inclusion of screen shots is an excellent use of screen space.
As-is, this is a solid article that I can't send to my team because of that one poor choice.
Here's why:
1) I don't know the personal histories of everyone I work with, and it's none of my business (they aren't expected to tell me).
2) Some people would be seriously distressed to have their boss or co-worker email them anything with remotely sexual content.
3) I can't reliably identify who these people are, and so I simply shouldn't send anything with sexual content to co-workers/employees.
Doesn't this same logic apply to just about everyone?
http://news.ycombinator.com/item?id=4280213
It quickly gained an additional 5 upvotes, and I was looking forward to the HN community's discussion, because I am about to revamp a certain site, and was looking to see what people thought about the specifics listed on that page.
It got to number 7 on the front page at about 11:10. Then it disappeared. It's now down between 900 and 950 or so. Yes, I went and checked.
I was disappointed - I always look forward to in-depth and/or illuminating discussions on HN - but more than that, I'm really curious to know why this item has been flagged so heavily, or, alternatively, buried by the moderators/admins.
So please, here it is again. You're (obviously) free to flag it again if you feel the need, but please, let me know why.
Thanks.
Added in edit: If you upvote this submission then you should probably go upvote grn's original: http://news.ycombinator.com/item?id=4280213 .
It's a minor detail, but an easy one to get right.
http://www.solipsys.co.uk/images/PasswordResetFlow.png
I agree that there's more going on, but this is beyond the original remit. The flowchart can be made seriously more complex by adding in the flow for the simple changing of a password, which needs similar attention to details.
Would you care to assist in producing such a chart? I have provide the DOT source for this image if you like.
Why is this NSFW? The picture does not show nude girls, actually the level of nudity is about what I can get to see in a public pool or on the street on a hot day. The picture certainly is suggestive and hints at something more. The article does not link to a porn site, load images from a porn site or even discuss a porn site. Is the hint that porn sites (or this specific porn site) exists sufficient for the NSFW marker?
I'm asking since this is certainly not considered NSFW for anyone I know - I can see "worse" in respectable TV documentaries which air in the regular evening programme. And I guess there's little chance of minors being around here either.
The other problem I have is, as you say, the image is suggestive. Imagine someone at your office glancing at your monitor whilst walking past. Would they bat an eyelid? I can't think of many workplaces outside of the porn industry where your professionalism wouldn't be called into question, however briefly.
Wow, that's interesting and surprising; where are you based, and what types of companies have you been working for?
I'm definitely haven't worked in lots of cultures, but have worked in offices in the UK, Poland, Belgium and Germany, for major companies and for small, and in each one of them, Questions Would Be Asked were I to have that image on my screen, and in most cases - but admittedly not all - my professionalism would be called into question.
In some work places I worked I'd be more concerned about coworkers/clients seing me reading a blog on paid time instead of coding than about that screencap.
Hehe, well that's a very different argument, and one that you and I agree on :)
I'm more curious about the fact that people could obviously be afraid that this might be a firing offense - and I can't think of a single place where this could be possibly used to construct a case.
That includes working at a very socially conservative workplace run by an evangelical christian and consulting for a government agency in a fairly corporate environment.
Some people might do a double take because of the subject matter of that image I suppose but nothing I would need to worry about, it's embedded in a technical article and it's PG-13 material.
I would worry far more about reading blogs than the PG-13 screencap that only really references the existence of porn sites.
I find it interesting and surprising that there are this many people with such a different experience. Are you working in very corporate environments?
That's not totally true. The screenshot fits perfectly well, because it not only shows the effect of a now know-to-be-known email, it also shows in which context this might be an issue.
Did you not see the word Porn in large, bold white letters against a black background with three young women just underneath it showing ample cleavage in suggestive poses?
There's a time and place for that, work is not one of them. And reading blogs (that are work related) is fine and even encouraged, but visiting sites with graphics like that is not.
[1] For me the "it's interesting and educating" trumps "there's naked women in suggestive poses" any time.
Some US employee avoid anything that could be used against them.
In some situation when they have to pick employee A or B for a bonus, a promotion or to be fired, and there is no clear difference between the two, the choice will rely on subjective and detail differences like beeing once caught with a porn page on its screen. Some employees could jump on such opportunities to smear their competitors.
Thus, since it was not required to use such type of screen capture to illustrate the problem, and it is not the usual type of content we read here, it make sense it was initially flagged and frown upon. Though I don't think fair to say the author is lacking judgment. To me it is the result of a culture difference. None good, none bad.
As a completely unrelated side note: I think that picture was an excellent choice. It's a graphic example that pretty much everyone can relate to. I'm not embarrassed by having that picture on my screen but I'd be embarrassed if my email address could be verified to belong to a registered porn site account [1]. It gets the authors point across so much better than showing an MSDN account.
[1] It doesn't. Don't bother looking ;)
But I am close enough to the US to see why people are nervous. Even acknowledging the existence of porn seems to be enough to make some Americans nervous.
I'm surprised at the number of people here making that many complaints about a PG-13 image, I honestly came here thinking I would see more people making fun of the NSFW warning for being completely reactionary.
Though: when a post from a personal blog makes it to the front page on HN, the author usually finds out fairly quickly, and will read the feedback here.
An easier fix is for users to randomize the email-addresses they use. For example, if you have a Google Mail account, you can receive messages at username+randomdata@gmail.com; if you use different random data for each site, nobody will able to probe for your account as described in the article, even on poorly-designed websites.
The advantage of this approach is that it's something that privacy-minded users can do themselves, without having to rely on website developers to get things right.
The reason it contains this is to illustrate that, if your services says "We have sent an email" vs. "User not found!" then it can leak the information that an account exists for a given email, which may perhaps be embarrassing given the content of the service -- especially if it tells you your User ID.
>> Everything i believe is important
In the spirit of completeness, you could expand the details of this paragraph:
>> What we want to do is create a unique token which can be sent in an email
There should also be a Time to Live on the token, and it should only be usable once. On the landing page for the link, the user needs to enter* their email address to avoid someone arriving at the url from a means other than the original email.
I feel like i should give an example there, let's say the attacker figured out your random id generation method, say it's a hash of a timestamp for simplicity, if they generate a few thousand links for around 6pm, they may get lucky).
* They don't actually need to enter it, UX guys would be having minor heart attacks at that suggestion, but they could choose their email address from a small table of plausible looking but made-up addresses, or something to that effect, e.g. Bank of America uses a photograph the user setup beforehand.
Just $.02 for this conversation :)
Bottom line: never give real answers to secret questions! I can change my password, I can change my credit card if it's compromised, I can even change my social security number if need be, but I can't change the name of my first pet or where I first met my spouse. Never give real answers to secret questions.
As expected, no one has said why it's been flagged.
Unexpected bonus, the original submission has gained enough points to be back on the front page despite having been flagged off it. So that's a win.
Added in edit: Now it's been flagged off again. There are times I really don't like the dynamic on HN.
It might be interesting for one of the sites that scrape HN to reverse engineer the decay factor for different type of stories and find out which stories are "disappeared" quicker.
The reset URL should also expire in about a day, so include and sign timestamp as well.
And the way we know this is good advice is that most large sites do it this way already.
You could include a sequence number in the token, but this, of course, involves a database write, which is what you were trying to avoid in the first place.
A better approach would be to store in the database the time that the user's password was last changed, and refuse to honor any reset tokens that are timestamped prior to that time.
if database.contains(email) or md5sum(secretSalt + email).startsWith('4'):
print "An email has been sent to this address with further instructions"
else:
print "Sorry, this address does not exist in our database"
Now one in 16 (pseudo-randomly selected) emails will appear to be in the database, whether or not it is really there.The attacker will still see the "confirmation" when they enter the victim's email address, but they cannot know if it was really in the database or if it was one of those random false positives.
- contained - false positive - not contained
Probably with the first two being close together and the latter two being close together.
Nice summary, sandstrom. Not as useful as reading the whole article, but much higher value-per-effort. ;)
edit: Not to mention that this chart still leaks the fact that joe@example.com is or is not a member of the site in question.
I'm off to invent distributed processing and BSD sockets.
Despite plenty of guidance to the contrary, the first point is really not where we want to be. The problem with doing this is that it means a persistent password – one you can go back with and use any time
If a developer can't work out how to implement a one-use password that forces a password change after login, they shouldn't be writing code.Sounds pretty far fetched to me.