300M Freely Downloadable Pwned Passwords
troyhunt.com
troyhunt.com
Currently, my way to generate a new password is this : `pwgen | md5sum`. And then, I use "lost password" everywhere (but for my mailbox, obviously), that is, the rare times my browser is not already prefilling the login form.
This makes me wonder why we don't just go with that : generate a random password for the user in registration form, allow the browser to save it. On the login form, check if fields are prefilled. If not, only display an email field and send an auth link as mail. User clicking it (once, and fast enough) is logged in.
You still have to remember your mailbox password, but that's the only one, quite akin the root password of a server.
Anyway, as far as password creation goes, here's another alternative based on more standard tools:
head -c NN /dev/urandom | base32 # or base64, or md5sum if you so prefer(note for people who may not know pwgen : it outputs 20 lines of 8 random passwords)
Unfortunately, it only does this if STDOUT is a TTY. If it's a pipe, then it outputs a single 8 character long password. So I'm afraid you've been generating rather low entropy passwords by piping the output of pwgen through md5sum.
You can confirm this by piping pwgen through cat: pwgen | cat
Most annoying trying to use "watch" or "less" and have the color work. "watch --color juju status --color", ahh. :-)
However, more recently I've grown skeptical. It leads to surprises, and if a person isn't familiar with the finer details of file descriptors, they may wonder why piping a program changes its behavior. I think this example of someone shooting himself in the foot with pwgen has convinced me never to do this again.
First, I used pwgen to generate temporary passwords for new unix users. I read the man page at that time, probably saw the mention about different behavior regarding output capabilities and thought : "nevermind, this is not my use case". Years later, when I decided to pipe it, I thought I already knew the program, and certainly not thought : "hey, let's check the man page again to see if the behavior may be altered on a pipe". Especially since I did not think it was a big deal, I was not putting that in a codebase, I just wanted to generate a random password to upload gifs or something.
And that's the interesting thing : should I have wanted to put it in a codebase, I would have read the man page again carefully.
All of this points to one conclusion : consistent defaults for end users are more important than sensible defaults for developers. You can expect developers to pay special attention, so that's ok to alter behavior for them through flags rather than detection of what stdout is plugged in.
If head behaved differently when piped, you would probably know about that because it's a common thing to do.
In terms of generating a strong password: most browsers have password generators (which I didn't even know about until recently). They aren't all enabled by default and they don't work on all forms, and not all browsers have them. Browsers also have supported user certificates [which are more secure than passwords, sort of] for like 15 years, but nobody ever uses them. The main reason (afaik) is how shit the browser's UX for them was, in combination with a burden of complexity - and of course you can't use them on random public devices.
I think we are close to reaching an authentication nirvana. If U2F came embedded in all new computing devices, and an open method of securely synchronizing all devices and service providers was used, we could effectively skip passwords, and rely almost entirely on backup codes for the few times they were needed. A lot of laptops and phones come with fingerprint scanners now. If those scanners were used as part of a U2F solution we would have a pretty solid authentication mechanism. (fingerprints are not foolproof, but IMHO they are about as secure as a password)
Pick a really good password for your password manager and remember that. Doesn't matter too much which password manager. If you're paranoid use KeePass, otherwise I personally use LastPass.
dd if=/dev/urandom bs=2048 count=1 | od -A none -l |
sha512sum
Or just generate a long password: pwgen 2048 1It's something we've come to embrace in the Linux world. Much faster than a single server and saves bandwidth at individual sites. Surprised this pragmatism hasn't reached the rest of you yet.
Even in your situation, if you're doing this for work and IT calls you up, you just tell them what you're doing. "I'm downloading a very large file over Bittorrent because it's 50% faster than the HTTP download and I'd like to do some work today. Is that all? Thanks. Bye."
In some places you'd get that call from IT just for downloading a large file. Getting a call doesn't mean you're doing something wrong, they're just checking it's you, not malware, and that it's for work. If they haven't already blocked it, have at it.
Who do we lobby to get them to fail their next PCI-DSS compliance test?
If that conversation starts off at 1000 characters you're fine, but more often then not it looks more like:
"Make the requirement 8-12 characters"
"12 is too short"
"fine make it longer, like"
"Ok" [18 char implementation]
"Hey, Bob in accounting says he uses 20 char passwords"
"Fine, bump it by another 50%"
[27 char implementation]
[no further internal complaints]
[Specs never updated or reviewed again]
Might happen in the next version release.
"No length restriction on passwords" is a common and valid report on HackerOne, because servers that do store passwords securely can be DoSed by someone providing a long password and forcing the server to hash it.
I doubt there's anyone who can make a strong argument that "64 characters isn't enough", and I doubt even intentionally computationally expensive password hashing is going to end up with significant resource usage with 64 or 128 character strings. I wouldn't want my shared hosting WordPress site with a password plugin to need to calculate the bcrypt hash of the entire text of War And Peace, bit I suspect the difference between bcrypting "password123" and a random 64 or 128 character string is insignificant enough to be ignored (but I've never tried benchmarking it, so I'm open to changing my mind here if anyone has links to benchmarks that show otherwise...)
Login attempts (should) get rate limited independently of password length limits, which makes the difference between hashing 8 characters and hashing 72 characters even less meaningful.
If you're working in an environment that has a FaaS (AWS Lambda, etc), it may be worthwhile to have this as an async function call so as not to block your primary application. Another option is to break authentication into its' own application, and have that return a signed auth token to your application.
There are lots of options.
Look, people on HN make a massive deal about passwords. One of my most shocking discoveries starting as a pentester was that "storing passwords in plaintext" would be a low-severity finding at best. Medium through critical vulns are reserved for findings that can own an app. That's how little password storage matters.
If you're relying on UPS preserving the secrecy of your 21-character master password that you're using across all your websites, you're doing it wrong. Yet the vast majority of users will do exactly that. The way to protect them is for critical services to use 2FA, which they do -- email, phone, insurance, etc all use 2FA or separate 4-digit passcodes now (USAA).
There have been so many password database leaks, yet the world moves forward. What is unacceptable is for Blue Cross to leak all your PII, yet the world moved on from that. CloudFlare leaked a huge amount of sensitive info. All of those matter way more than some password leaks.
If someone is going to target you, here's the most likely method: https://news.ycombinator.com/item?id=14919845
I suspect that's because you're viewing the situation as a pentester not a user. A plaintext password (on its own) doesn't do a pentester much good until they've already gained control of the system. However, once someone has control of the system then plaintext passwords are a threat to users because a lot of people are vulnerable to common password reuse.
(This application has had three failed rewrite/replace projects over the last 10 years. They're now onto their fourth one, and they're running 32 months over deadline on a "2 year" project timeline...)
With a limitation of 27 characters it's not possible to make my password be "if i forget this i'm in a lot of trouble".
At least one bank says the pin is "for this device" but then accepts it on new iOS devices too (without ever prompting for the "real" password on the new device).
(I occasionally worry that my phone's bitcoin wallet does not do this. There is occasionally a large enough balance in there that I'd really like to have to touchid to open it and transfer those bitcoin out... Not quite worried enough to investigate whether other wallet options do it, but sometimes I hand my phone to someone to show them a pic and think "Do I _really_ know this person well enough to trust they won't poke around my phone and try to distract me enough to swipe some bitcoin???")
Don't Fuck With Paste: https://chrome.google.com/webstore/detail/dont-fuck-with-pas...
(And yeah, the "internet random" here has a github repo with the code, and the file that does this is an easily auditable 16 lines of javascript, so props to him for that. But it's still got the recently exploited attack vector that he or an attacker who takes over his account could push malicious updates to the extension, like the webdev extension from earlier this week...)
I then entered {DELAY 3000}{PASSWORD}
Now I can log in to a full screen game that doesn't allow pasting (I type in the username by hand first). If I'm alt-tabbed out of the game with KeePass in focus, the three second delay is enough time to go from triggering auto-type to restoring the full screen game window and clicking on the password field to give it focus. I was unable to trigger KeePass autotype with the game already full screen.
If the account you're trying to log in to also has a long random username that needs to be auto-typed, you'll probably want a sequence like "{DELAY 3000}{USERNAME}{TAB}{PASSWORD}". Or, if the form doesn't allow you to tab from the username to the password field, you could use "{DELAY 3000}{USERNAME}{DELAY 3000}{PASSWORD}" and you can click on the password field during the second delay.
The password you create here can be used to access Online, Mobile and Telephone Banking.
All passwords must be six characters in length. Special characters (eg. *, %, $, etc) will not be accepted.It may sound archaic but you have to raise the visibility if you're concerned about changing the banks behavior.
>Do not send any password you actively us to a third-party service - even this one.
So I can only test password that I am not using (and by extension that I am not going to use in the future).
>oh no - pwned!
>This password has previously appeared in a data breach and should never be used. If you've ever used it anywhere before, change it immediately!
If I cannot (shouldn't) submit any password I am actively using, what does it matter if I used it before? Now I already changed it.
The gp makes a good point, but that's also why you can submit the `sha1($your_password)` instead. The only question is why did Troy allow un-hashed passwords to be submitted.
I mean, here is the SHA1 of my password (not really):
d012f68144ed0f121d3cc330a17eec528c2e7d59
This site:
https://hashkiller.co.uk/sha1-decrypter.aspx
>We have a total of just over 312.072 billion unique decrypted SHA1 hashes since August 2007.
Took exactly 221 ms to reverse it to "pippo".
Troy mentions some arguments against torrents, but it is better to have a authoritative torrent than none, imo.
I'd probably just shove the passwords into a database, limiting the index prefix to the first X characters to reduce index size.
I distributing to a general audience, 0.5GB and 10GB isn't that much of a difference, and most people are more equipped for handling lists of strings than for handling bloom filters.
Can it? I think of a bloom filter as similar to a lossy compression scheme and wouldn't expect it to be further compressible to any significant extent using a general purpose lossless scheme. Similar to how general purpose compressors generally don't do very well with mp3s or jpgs.
The nature of the data means it can never really be "perfect" anyway (there are certainly some password breaches that exist but aren't included in the list), so massively reducing the resources required in exchange for a bit of artificial error seems pretty reasonable to me.
For false positive rate of 2^-9 (0.00195) that's 447 MB, which is slightly less than an optimal bloom filter for the same number of items, and lookups will be considerably faster.
Construction time and the fact that you can't add to it without rebuilding the whole thing are the downsides. But given the application I don't think they matter much.
http://sux4j.di.unimi.it/docs/it/unimi/dsi/sux4j/mph/Minimal...
> I'm envisaging more tech-savvy people using this service to demonstrate a point to friends, relatives and co-workers: "you see, this password has been breached before, don't use it!"
But I can't be the only one whose family would be baffled by the term "pwned". I wish it said something like "Your password has been hacked!" which we all know not to be technically correct but would resonate a lot more.
Wait, so I test my password to see if it's "good" and now you have a copy of a password I will be using. Am I just being paranoid?
Maybe a malicious copy of the website exists at lots of LevenshteinDist=1 domains. Accidentally typo the domain and get pwned, thinking you are submitting it to an ethical security researcher's tool, but actually getting phished.
$ sha1sum
SooperSekretPassw0rd^D
SooperSekretPassw0rddc0d3504b259a92dce59b850969601d12c06a75f -$password = "foobar"
([Security.Cryptography.SHA1CryptoServiceProvider]::Create().ComputeHash([Text.Encoding]::ASCII.GetBytes($password)) | %{'{0:x2}' -f $_}) -join ""
"foobar" | Out-File password.txt -Encoding ASCII -NoNewline
Get-FileHash password.txt -Algorithm SHA1
del password.txt
That's readable and intuitive, but the downside is it puts your password in a file. /tmp$ echo "p@55w0rd" | sha1sum
8633c4a8b38a8826132414d8861af7b6a8371976 -
This is a different value from the one given in the blog post: "ce0b2b771f7d468c0141918daea704e0e5ad45db".The python sha-1 hexdigest comes out right, though:
In [13]: import sha
In [14]: sha.new('p@55w0rd').hexdigest()
Out[14]: 'ce0b2b771f7d468c0141918daea704e0e5ad45db'
In case anyone else has passwords they want to check, this will binary-search them: https://gist.github.com/coventry/5df7885fb0d5caeabb39fcd0e2b...For example, if I have a site that requires passwords to be at least 10 chars long, I don't need any of the data for breached passwords that are shorter than 10 characters. People can't possibly use them anyway, so that's probably a huge chunk of the data that's completely useless to be storing and checking.
https://gist.github.com/roycewilliams/b1de2afbfe5cb71bea16c9...
Regardless of composition, the top 12 lengths are:
8: 32% (102260862)
10: 14% (45084047)
9: 13% (41525797)
7: 10% (33632055)
6: 06% (20211176)
11: 05% (18275968)
12: 04% (14052958)
15: 02% (8291459)
13: 02% (8042452)
14: 01% (6321198)
16: 01% (4201765)
5: 00% (3054291)
In other words, requiring a minimum length of 12 would make 80% of the passwords in the corpus inapplicable.... and the top 12 masks are:
?l?l?l?l?l?l?l?l,47823614
?l?l?l?l?l?l?d?d,7005728
?d?d?d?d?d?d?d?d,6212778
?l?l?l?l?l?l?l?l?l?l,6023602
?l?l?l?l?l?l?l?l?l,5379482
?l?l?l?l?l?l?l?d?d,5169013
?l?l?l?l?l?l?l?l?d?d,5090400
?l?l?l?l?l?l?l,4998896
?d?d?d?d?d?d?d,4798329
?l?l?l?l?d?d?d?d,4798124
?d?d?d?d?d?d?d?d?d?d,4754401
?l?l?l?l?l?l?d?d?d?d,4377841
Almost 48 million of them are 8 lower-case characters.* And to be clear, "cracked" is an overstatement. Many of his sources are public. Simply using those sources as wordlists makes "cracking" these like shooting fish in a barrel.
Are you planning to make a blog post or anything "final" with the info you find out, or will you just keep updating those gists?
With a tool like hashcat, a modern GPU or two, and some publicly available wordlists, you can get the vast majority of them without breaking a sweat.
In other words: there is almost no value in hashing them with SHA1.
"It goes without saying (although I say it anyway on that page), but don't enter a password you currently use into any third-party service like this! I don't explicitly log them and I'm a trustworthy guy but yeah, don't."
Safe: probably. Good practice: no.
Well, I better get prepped to fight the Estatis Inc. Retaliatory Creature to retain full ownership of my soul. Thanks for the heads up.
It is much, much safer to download the data and search for your passwords from inside the local copy of the data, and that's exactly what I'm gonna do later tonight.
Guessing it was in MySpace..
Ironically I used another password for sites I trusted less and that one isn't in there.
Also, I'm genuinely curious as to why SHA-1 is used and not SHA-256. Surely the one-time additional cost of using SHA-256 would've been negligible for Troy? If at some point somebody manages to do preimage attacks on SHA-1, I have to assume my password is broken if I've submitted its hash to his API. Although I guess you'd have to actually be able to enumerate preimages, preferably from small to big. Still, I don't understand why Troy doesn't account for the possibility by using a hash function widely considered to be stronger.
Go here and type the character 'a':
BTW, I upvoted your comment. It made me laugh. Point granted.
If someone's website forces me to set up an account without me already being convinced there's enough benefit to me in return for my personal information, they're likely to get a signup for test@example.com with password "foo". (And if they then respond with "please click the confirmation link in the email we just sent", I'll sign up again with $sitename@$spare-domain-I-own.com and pick up the link from the spam filtered catchall account).
I probably do this at least once a month when there's hints something useful in a web forum I want to read, but I'm not (yet or ever) convinced I'll ever become a member of that forum's community...
The second google result for rainbow tables lets me download software and tables to efficiently reverse any sha1 whos plaintext fits [a-zA-Z0-9]{1,9} or [a-z0-9]{1,10}. That's likely the majority of passwords an attacker would observe
> Each of the 306 million passwords is being provided as a SHA1 hash.
That's it? Without any salting? This would make it trivial to recover the plain text using rainbow tables.
It's not like everyone who is curious cant go findabout stuff like this doesn't have the Rockyou dump already, and there's easily enough links to start your own list here: https://www.google.com/search?q=password+lists
I have no experience with Scaleway on this, but based on what I've heard about them in the past, I imagine their policy is roughly the same.
But 306 million passwords have already been exposed by data breaches at other sites. And users have a tendency to reuse passwords across multiple sites. Just because your website wasn't hacked doesn't mean an attacker can't go look up one of your users in someone else's data breach and try the same password on your site.