Pwned Passwords in Practice: Real World Examples of Blocking the Worst Passwords
troyhunt.com
troyhunt.com
But reading through the API docs [0] it shows that API has no rate limit. Impressive.
[0] https://haveibeenpwned.com/API/v2#SearchingPwnedPasswordsByP...
Not saying is a bad thing but don't assume something is for pure altruism because not many things are.
Complex passwords are quite useful if the server gets hacked and someone walks away with the (salted) password hashes. Against brute-forcing passwords at the login screen of an application they don't add much value, other than making it the user quite hard to remember what the password for this particular site could be...
Theoretically, if you block a user ID after say 5 or so invalid logins, almost any bad password from the Have I been P0wned list will prevent you from being hacked. The chance that you pick exactly that password from the 1-million or so list is quite minimal.
So with that in mind, wouldn't this service be something for website owners that don't know how to properly secure the information they control?
1. If one of your passwords is suddenly rejected, it may be a great moment to refer you to HIBP.
2. Traditional password complexity estimations may be overestimating your passwords complexity, e.g. the pass phrase "My house is blue" is fairly long and will likely be flagged as complex enough. But it's within the realm of a password-phrase aware tool.
From experience, most attacks we see now are credential stuffing attacks rather than pure brute force attacks using something like Sentry MBA, with a huge number of IP addresses (the last attack we saw was using over 6 million IP addresses). So throttling sign in attempts at the IP level is almost useless as is throttling at the email level, as the attacker can attempt at least 6 million known email/password combinations to see if those accounts exist on your site.
The only real defence against that is all your users using 2 factor, or creating a psuedo 2nd factor (email them if the attempt is from an unrecognised IP).
Edit: Of course the other helpful defence is to ensure your users aren't reusing passwords, which is where Pwned Passwords comes in.
Disclaimer: work for Kogan who is mentioned in TFA.
What do you mean by "end up succeeding"? Most requests successfully authenticated? On the first try? Second try? Tenth try? Hundredth try?
(I'm not trying to doubt the utility of pwnedpassword validation; just hoping you can help me understand the threat you're facing and why IP rate limiting didn't help much. Thanks.)
Lets say you have IP throttling/rate limiting. And you have it set to an extremely conservative limit - 1 sign in attempt every hour. This is great for the brute force threat - 24 passwords a day can be attempted by 1 IP. Infeasible for any brute forcing.
But now lets say the attacker has access to a botnet with 6 million unique IP addresses (not theoretical - see my comment above).
Now for each of those 6 million IPs they can try 24 passwords a day - i.e. 144 million attempts a day without ever triggering the throttle.
Bear in mind also that they aren't just trying random passwords for an account - they have a compiled/combined breach list of known account/password combinations from other breaches. So they can attempt 144 million known combinations a day. Without hitting any throttles (this is what the parent above means by "end up succeding").
What percentage of your users reuse passwords and have been exposed to at least one breach? I would suggest it's quite a high value. How long do you think it will take a credential stuffing attack to identify those accounts on your site when they can try 100's of millions of combinations a day?
This is the threat vector.
Why?
If he comes to my site and signs up with that password, an "evil person" doesn't need 5 guesses to get into his account - they just need one, because they already have it.
If, however, I check Jimmy's password when he registers, and block him from using it: (1) I keep him from immediately losing control of his account on my service, and (2) I provide Jimmy with the knowledge that his favorite password was leaked and he needs to do something about it.
Yet, this has become almost like a public good over the years, and providing it at no (extra) cost than fundamental costs of accessing it is appreciable.
That seems way too low.
For context: If you take just dictionary words from the world's 5 most popular languages, you'd have more than 0.5 million words.
https://haveibeenpwned.com/API/v2#SearchingPwnedPasswordsByR...
With a good cache, thatd save some bandwidth.
Maybe that's wishful thinking. I can't imagine checking new passwords more than a few dozen times per second at the most. Bigger sites probably just write their own password integrity tools.
Most places just enforce byzantine password requirements, 13 digits, must have ~, uppercase and a palindrome prime integer in it.
Obligatory password XKCD, think of the children. https://xkcd.com/936/
I use a memory trick to have very strong passwords, but most people probably wouldn't be willing to invest the effort.
Someday there will be a better way.
Shown in API docs under "Pwned Passwords overview" and links to here: https://www.troyhunt.com/enhancing-pwned-passwords-privacy-b...
I made a start on a Ruby port too: https://github.com/Freaky/ruby-gcs - I have vague plans to finish it off and write a Rodauth (http://rodauth.jeremyevans.net/) plugin for it.
I wrote a devise extension thats essentially a one liner to add this check on signup (and a small code block to add on signin)
https://github.com/michaelbanfield/devise-pwned_password
Nowadays most of the logic is encapsulated in the pwned gem
https://github.com/philnash/pwned
Which is a good choice if you arent using devise.
Edit: On third thought, a bloom filter for 502M entries and a false positive rate of 0.1% ends up as a 800MiB large filter. Binary-searching the whole dump is surely faster.
With that sort of FP rate it's not really much use beyond filtering API calls. I'd suggest 2GB[1] as a more sensible minimum. A compressed filter can get this down somewhat.
> Binary-searching the whole dump is surely faster.
Not really. log2(500M) is ~29, k for a suitably sized bloom filter's only 23. Interpolation search can get you a result in more like 10 seeks, but a bucketed bloom filter can get your lookup down to a single read.
Having spent a fair bit of time faffing about with this stuff I ended up settling[2][3] on Golomb compressed sets[4], which can get the full list with a 1-in-10 million FP rate into 1.5GB.
[1]: https://hur.st/bloomfilter/?n=500M&p=1.0E-7 [2]: https://github.com/Freaky/gcstool [3]: https://github.com/Freaky/ruby-gcs [4]: http://giovanni.bajo.it/post/47119962313/golomb-coded-sets-s...
Yeah, it's just k single-bit lookups - ideally you do something to get them into clusters, like dividing the database into sub-filters, so you're doing random lookups into, say, a 32KB chunk instead of a whole 2GB filter.
> Did you look into Cuckoo filters as well?
Cuckoo filters look like an interesting alternative and looking at them more closely is on the to-do. I don't think they'd have any significant space savings, though - they're similarly about 75% the size of the equivalent bloom filter. Maybe they'd be faster for lookups?
I'd also be interested in playing with matrix filters[1], which supposedly get close to the theoretical limits for these sorts of structures. Implementing them seems rather more involved, sadly - particularly given the only reference I can find is a fairly inscrutable CS paper. Show us the code damnit.
https://haveibeenpwned.com/Passwords
Its a losing battle trying to add byzantine rules to prevent users doing things like using their normal password * 2, so its probably a reasonable check to add.
Still, I'd bet the vast majority of the bad passes are <16, seems a heck of a waste of energy and bandwidth to check my user's passphrases against (guesstimating) 0.05% of the corpus.
OP really cares about his privacy/security/password, but then he uses a secondary system to store it?
Is this the best way to do this? Break into the secondary system and everything is available. Keyloggers, eyes, stolen computers, all have the possibility of everything available.
I considered other solutions like written + put in a lock box in a bank, but thats really inaccessable.
Modern password management services are incredibly secure, with client-side encryption of your secrets, among other protective mechanisms (to guard against keyloggers, stolen computers, ...)
There is still a risk, because you're trusting third-party software (which in some cases is closed-source – including 1Password), but for most people that is a much lower risk profile than if they were managing passwords themselves.
For specifics on 1password's security, check out https://1password.com/security/
There lots of password managers that have had gaping vulnerabilities. From memory I'm sure I've seen LastPass vulnerabilities top HN just a while ago... yeah probably this one: https://www.bankinfosecurity.com/lastpass-patches-password-m...
FWIW I'm not a fan of keeping them in some cloud service dedicated for storing passwords (1PW's subscription offering, lastpass) because if someone is going to break into servers and steal passwords it's going to be from those fine people.
A middle ground is to use something like KeePass, where the encrypted password file is stored where you want it to be stored (ie your hard drive), where you can manage how you want to distribute that file between places you need it.
I still have to rely on the mobile app which isn't open source, but it is the recommended app from the main developer.
On the security spectrum, for the vast majority of people, a password manager is a step up in how safe their passwords are.
What could possibly go wrong?
This works by locally hashing your password, then sending only the first 5 hex characters of the hash to the server. The server sends back all matching hashes of bad passwords it has on file. Typically this returns a few hundred hits. The local client (probably some piece of Javascript) checks the hits against the password hashes returned. If there's a match, the password is in the database of bad passwords. This supposedly protects the password if communications with the checker are intercepted.
But does it? If an attacker can see those first 5 hex characters, they too can get the list of hashes of matching passwords. There are only a few hundred of them, and they're hashes, not the actual passwords. One of those is the hash of the user's password. Now they know what hashes to try.
An attacker presumably has a big database of likely passwords to try. So they can create a database of hashes locally. The hashing algorithm is known to the client, after all. So now they have a few hundred passwords to try for a break-in. Try those over the next few days, and they're in.
Is it really that bad, or am I misunderstanding something here?
Why? Your hash might not be in the list. There are 16^35 hashes starting with these 5 characters.
But how do we trust the Okta chrome extension not to post all credentials to the developer? Even if its doing nothing shady now, in some future update?
At least with this one, you can audit the source code and build it for yourself[0]
But... if we'd wanted to be nasty we could have sent everything to servers we controlled.
The src link stopped working when gitlab.io arrived I think https://github.com/markolson/js-sqlite-map-thing/
I'm pretty sure this is safe, but if there's a way to defer sending an HTTP request to after the page being closed...
And with the ease of using botnets now, they can get access to a very large number of IPs to distribute the attack, so you can't reliably throttle at IP level either.
Now imagine having a site with a few million accounts and 0,1% of them mistyping the password every now and then.
I think I’ve had this happen to a few old accounts of mine that reused a simple password that had been leaked. (Of course I can’t really verify how the hackers did it, but it seems likely to me)
Sort of like how credit card providers could make things much more secure but choose not to because of fear of reducing transactions.
And, if you do a hard lock-out, that’s an easy way for an attacker to DoS your entire user base in short order.
No seriously, Troy's site is awesome, and spreading FUD is not doing them any service.
It's a very useful service - and free. No way am I going to manually download every single password dump that appears online and search it for my details.
That said, the domain name is (IMO) needlessly clever and does not help instill trust in folks who are not already familiar with it. I recommend it to friends and family, but I cringe a bit whenever I have to post that URL into an email to, say, my aunt who is a retired librarian.
Treating your email address as a protected piece of information is a laudable goal, but seemingly impractical in the real world.