A decade of Have I Been Pwned
troyhunt.com
troyhunt.com
I don't have anything to back this up, but my guess is that the vast majority of compromised user accounts comes from credential stuffing/password re-use. It's really surprising to me when I hear that huge companies don't do this check.[2] It's simple, easy, takes about a day to set up.
If you're a young CTO or early-stage engineer working on a web app and have never been targeted with a credential stuffing attack, let me tell you: It's coming! It's just a matter of time before it's 1AM and your phone blows up; your site is getting hammered; you think it's DDOS, but then realize most of the hits are on your login page, then realize that and then realize with a horrible feeling that some % of those hits are getting through the login page. You'll be up all night dealing with it, and then you have to make breach notifications, and that really sucks.
Troy Hunt's free database will save you that heartache (probably). Just do it.
1. https://cheatsheetseries.owasp.org/cheatsheets/Credential_St...
2. Like 23andMe. https://news.ycombinator.com/item?id=37794379
The alternative is the exact same scenario, except that the percentage is several orders of magnitude lower, right?
The small subset of your users that explicitly opted-out of 2-factor authentication (if you allow that) and who try to choose "Password1!" with a second exclamation point when your site said "Error, your password has seen 83,000 times in password dumps, please use a unique password" will still get hacked.
Or is your expectation that no one will attack every user on your webapp with a credential stuffing attempt if they see that the probability of success is 0.001% instead of 1%?
Remember that "identity theft" is marketing fluff. In a credential stuffing attack your business is the victim of fraud.
Your numbers literally turns a scenario where 200,000 accounts are hacked into one where 200 are exposed. Or one where 30 hacked accounts turn into 0 hacked accounts.
There is a point where a difference in quantity becomes a difference in quality. I far prefer the latter scenarios.
The number of times I’ve seen DEVELOPERS neglect to implement materially useful security measures because “they’re not technically perfect!” Is astounding.
One of the employees was apoplectic at the actions of the sys admin and had accused him of violating her privacy by doing this. While I do not recall which party initiated legal action against the sys admin that led to his arrest (i.e. the employee or the company), the bottom line of the story was that the FBI employee (and, by extention, whichever judge was involved in adjudication the case) considered the act of a sys admin accessing password hashes placed under his care to be a criminal breach of privacy regardless of his intent being to improve his company's security against password stuffing attacks.
Assuming the FBI employee didn't just make the whole thing up (which I have no reason to believe - there are a lot of tech-stupid judges and, especially a decade ago, tech-stupid FBI employees), it might be prudent to pass this by your legal team before checking for password hashes for your employees being in haveibeenpwned.
That said, there are still tremendous gaps yet to be bridged with the understanding of many procecutors and lawyers as well as weird applications of the law that aren't intuitive to people whose life is technology.
For example (and I caveat this with IANAL): Did you know the physical medium you get Internet to your house determines what laws and processes the government can use to monitor your Internet traffic?
That I did know, only because I was dumb enough to hitch my wagon to Comcast/Xfinity as a headend tech for years. Just affirmed the idea that all ISPs should be community owned.
I did not! Do you have / know a good explanation of the details?
Other means of Internet getting to your residence is covered by Title 3 of the ECPA which, historically, Feds have played fast and loose with getting data from.
Like https://www.eff.org/deeplinks/2016/07/ever-use-someone-elses... > Last week, the Ninth Circuit Court of Appeals, in a case called United States v. Nosal, held 2-1 that using someone else’s password, even with their knowledge and permission, is a federal criminal offense.
Also, the courts only just legalized white hacking last year. Before that violating the terms of service was also potentially a federal crime. https://www.spiceworks.com/it-security/security-general/news...
"In July 1995, Schwartz was prosecuted in the case of State of Oregon vs. Randal Schwartz, which dealt with compromised computer security during his time as a system administrator for Intel. In the process of performing penetration testing, he cracked a number of passwords on Intel's systems. Schwartz was originally convicted on three felony counts, with one reduced to a misdemeanor, but on February 1, 2007, his arrest and conviction records were sealed through an official expungement, and he is legally no longer a felon." -- https://en.wikipedia.org/wiki/Randal_L._Schwartz
"Rather ill-advisedly, the Perl-programming guru (who's written several books on the subject) tried to prove his worth by running a password cracking package after he'd left in order to produce evidence that security practices had deteriorated since his departure. Instead of re-hiring Schwartz, as he hoped, Intel called in the police and he was charged with hacking offences."
https://www.theregister.com/2007/03/05/intel_hacker_charges_...
Also you are going to spend a long time being arrested before the appeal goes out.
"In California, a criminal appeal can take several months to several years. The length of time depends on the complexity of the case and how quickly it moves through the appeals process."
If I'm not mistaken it's all done using cryptographic schemes that leak neither the password nor the hash.
The Open Web Application Security Project's Application Security Verification Standard recommends that you do a hashed password check [2].
For bigger companies, sure, go talk to legal, but for young startups, my feeling is it's not worth the $200 or whatever your counsel will charge to say it's ok. I personally did not ask anyone (am cto), I just added the check.
1. https://twitter.com/troyhunt/status/1674132801837477888
2. See OWASP ASVS 4.0 2.1.7 https://github.com/OWASP/ASVS/blob/master/4.0/en/0x11-V2-Aut...
That said I struggle to believe the sys admin had competent representation.
Probably closer to $2000 than $200, but paying for an opinion is truthful, helpful and useful.
Kinda sucks that it's necessary
Parent commenter never mentioned anything about comparing stored password hashes. What you do is block bad passwords at password set time by hashing the prospective password and comparing with HIBP. A prospective password you haven't accepted or stored or transmitted off the application server - common sense says that's not a privacy violation - and many giant companies including my employer do this routinely.
[Edit] Oh yea I remember HIBP has an online API. Don't use this. Take the HIBP dumps that they make freely available and compare locally. If not for reasons of privacy, for reasons of simplicity and removing an unnecessary external business/legal/software dependency.
You can set a flag on login to use the password in memory rather than stored.
By the way, I don’t feel paranoid to flag bad passwords on login (perhaps triggering an email OTP and forcing a password reset), personally. I responded to this thread because a commenter made an unfounded implication about using HIBP data to reduce vulnerability to credential stuffing.
That's not the greatest advice IMO. The API gets updated data more frequently, doesn't require that you transmit the password or a useable hashed form, and it's dead simple to consume. I'd argue that it's more effort to maintain an internal store and synchronization infrastructure, and you're less likely to accidentally breach anonymity and leak a weak hash by using the API than you are rolling your own query against the raw data.
It's also used by hundreds of bigcorps and government agencies who have way more pedantic lawyers than you're likely to have. If they couldn't find a good reason not to use it I doubt yours will.
Just as many arguments can be made for an offline check. Or against an online check. From added latency via required uptime to added dependencies.
My point being: no. "It depends"
I know a company that started doing quarterly brute-forcing of passwords as a security check and the reaction to finding out that your password is not strong enough is....all sorts of emotions.
If you have a 10-12 character password that may have been strong at one point but now is not and your IT team is informing you, you're reaction is NEVER, oh thank you for helping me out. It's not stupidity, it's human nature to feel attacked.
Like imagine how many failed attempts must've happened for a 12 character password to get bruteforced. Alarms should have been raised way before it became an issue.
also, at one point it was popular to use l33t speak for passwords so there are many crappy 12+ char l33t passwords floating around that are trivial to guess, no brute forcing required.
That's what haveibeenpwned.com is about. It tells you if your email is in one of these database lists out in the wild. If it is, assume your password will eventually be discovered.
But there's more than just the issue of discovering the passowrd itself.
What about the issue of discovering that a particular password hash comes from an employee at a certain company.
As I understand it, Tory Hunt downloads dumps of stolen passwords. He does not share the dumps. Instead he collects queries, like a search engine. Until people start sending him queries of hashes to check he does not necessarily know the locations of the people whose passwords were stolen.
However if he gets a series of hashes sent from some IP address belonging to a perticular corporation, then argubaly he now knows these are likely to be passwords belonging to employees at that corporation.
https://www.troyhunt.com/understanding-have-i-been-pwneds-us...
It sounds like it was made up, should not be so hard to find the verdict.
This is a good way to disincentivize prosocial behavior.
The problem can be "who defines good deeds?" There are so many things which seem good when presented one way, but can be harmful when viewed another way. Obviously, as presented above this seems like "an obvious good", but context matters, snd clearly you don't get the whole context from a one paragraph summary.
Ultimately we have civil structures (government at every level) that tries to codify "good" and "bad". Life is seldom that clean though, so inevitably every regulation and law is good for some bad for others.
So, to answer your question, because "good" and "prosocial" are not universally true.
They're often now upset they've been called to task so it's just hard all around.
(Non sarcastic), why would you feel bad for users using 1234 as their passwords? Unless your website is aimed at vulnerable people, I consider this to be their responsibility.
As other comments have said these users will probably go the easiest route (1234websitename) to fix the error.
Any restriction you put on your password field reduces entropy, and safety for everyone (even if marginally so).
Breach notification etc legislation in some jurisdictions will also require that you report successful widespread credential stuffing.
Even AWS with their “shared responsibility model” works with GitHub etc to ensure that programmatic access credentials aren’t accidentally exposed via public repositories. This isn’t credential stuffing, but it’s a blindingly accurate demonstration of the fact that drawing a line in the sand and saying “users, work it out from here!” and attempting to wash your hands of the situation is nothing more than the ill-informed pipe dream of someone that’s never had to deal with this stuff in reality.
1234websitename is objectively better than 1234.
I'll go with NIST on this one (yes, and have a minimum length too):
> When processing requests to establish and change memorized secrets, verifiers SHALL compare the prospective secrets against a list that contains values known to be commonly-used, expected, or compromised... If the chosen secret is found in the list, the CSP or verifier SHALL advise the subscriber that they need to select a different secret, SHALL provide the reason for rejection, and SHALL require the subscriber to choose a different value.
For example, let's say Tumblr was hacked and with it my password `hunter2`. Tumbler used some naive HMAC-MD5 method with a salt, but my site uses argon2 with (obviously) a different salt. Even though my password is the same (`hunter2`) the resulting hashed passwords will be different. How is this any effective preventing credential stuffing?
Sorry, I still don't understand the procedure you mentioned and I'm genuinely curious.
So, the procedure you need to implement is, on login/registration/pw reset, you SHA-1 hash the user's unhashed password and do a indexed lookup on your copy of HIBP's database. Or if you don't want to maintain that copy, you can use HIBP's API to do something similar.
My understanding this isn't true. These leaks are often just the password hashes.
I do think checking against the HIBP DB is a good call too, but it doesn’t stop this attack overly well, rate limiting is a much better way to stop it.
But there's also "stuffing" with known breached username+password combinations – in which case it still helps, but I don't think as much? In the latter the attack is much more likely to succeed and there's a much smaller number of values being attempted, so the threshold of detection + blocking would have to be much lower...
If you're working on a greenfield login/auth, please don't accept and store passwords in a database! Setup social OAuth, SSO, or magic link emails and make it someone else's problem.
You don't want to end up with a naive implementation of OAuth2 (like some big names had recently) which fails to check the audience parameter, and therefore lets anyone other service using the same SSO gain access to your users' accounts.
Recent HN post on this - https://news.ycombinator.com/item?id=38009291
I use unique email addresses (breach canaries) on every website to detect when sites leak my data. When I tried to search for my domain results with a previous domain ownership verification, I got hit with this error: "In order to search a domain with any more than 10 breached accounts on it, you need a sufficiently sized subscription"
To make matters worse, Troy includes public data compilations as 'breaches' which artificially inflates counts for the breached accounts quota. For example, when a compilation of public contact details scraped from GitHub leaked, Troy counted that as a breach. I explicitly listed my email address as public.
I'd be willing to pay $5-12/year. These rates are outrageous for such a low-overhead service.
I have no idea if HIBP is a cash cow for Troy. It may be, given these prices, but I don't know much about his other sources of income.
HIBP didn't start out as a cash-grab, but it is one now. Troy could have chosen to price it reasonably to cover the costs of the service. This pricing is clearly taking advantage of HIBP's popularity as the de-facto breach list site.
I would bet you that over 99.99 percent of HIBP's users do not pay for the service. Troy's time has value, so working on a service that provides no income is not really something you can expect a person to do. Troy decided to create an enterprise subscription service to get a bit of revenue from something he's created. It's not cheap, but it's not something you're meant to buy unless you're a company looking to monitor your employee email addresses. This service is pretty cheap in that regard, actually.
I really do not understand why you feel that you're being ripped off here. This is just a lack of product-user fit, his pricing structure simply doesn't work for you because you use email canaries. But for a company with 100 people, this pricing is entirely reasonable, if not something incredibly cheap.
What you're paying for is everyone who doesn't pay for the service, the time he takes to add new breaches to the service and the time he takes to develop the service.
Just because their pricing model doesn't fit you doesn't mean it's a cash grab. Is this too hard to understand?
EDIT: Also, I assumed this was a common understanding, but product pricing is based on the value it gives to the person buying. For a company of 100 people, do you think paying $160/yr is worth breach monitoring? I think for any IT department, this would be a no-brainer.
It should actually be considered best practice.
Email canaries clearly flunk the second test.
(UPDATE): I’ve posted a suggestion to the UserVoice community, which it appears Troy actively monitors.
If the several (dozens?) of us with this use case upvote it, it may catch his attention.
https://haveibeenpwned.uservoice.com/forums/275398-general/s...
Basically he suggested doing a monthly subscription for just one month periodically as a way to reduce the cost.
Another option is to read the notification email, and if it’s for Acme Corp, to remember that the associated email must be acme.com@mydomain.com, and then manually check that.
That's a price of its own.
That is honestly pretty hilarious of a side effect of media fame!
Years ago we had friends, a couple in which the wife was pregnant. They were actually a bit embarrassed that “everyone will know that we ‘did it’”. A level of squeamishishness I could not have imagined!
Most people would be more embarrassed if the wife is clearly pregnant and nobody thinks they “did it”.
Although, maybe it could be a database of facial recognition hashes from revenge porn, and you can upload a similar hash of your own face to see if you're online somewhere?
I sometimes think what I would have done had I never read his posts about checking without transmitting real PII.
I'd also like to call out the one who Troy says suggested him [1], Junade Ali who goes into more details about this in his post about it [2]
Not because Junade would have invented it (apparently that was Pierangela Samarati, Latanya Sweeney and Tore Dalenius. [3]) but because his blog post on it is a really great explainer of it using concepts software developers are familiar with.
[1] https://www.troyhunt.com/ive-just-launched-pwned-passwords-v... [2] https://blog.cloudflare.com/validating-leaked-passwords-with... [3] https://en.wikipedia.org/wiki/K-anonymity
I assume most get blocked by spam filters. I've only noticed them when they get past SPI/DKIM filtering and I have to train more. They seem pretty clever.
I appreciate the service and enjoy Troy Hunt's posts. HIBP is great.
I can’t imagine how it would feel if I didn’t use a password manager and couldn’t quickly see where was that password used instead of wreaking my brain trying to remember!
I got one or two of these emails, and couldn't for the life of me guess on what sites I had used the (pretty weak) passwords.
On the other hand, I feel like SpyCloud does not get enough credit for having a dataset 30x bigger and working directly with companies to actually mitigate credential reuse. If you've ever been prompted at login to a major website or received an email asking you to reset a password because it was used in multiple places, there is a good chance SpyCloud is behind it.
Not sure if it is your intent, but this implies that HIBP does not work directly with companies to mitigate credential re-use. It does. Some examples would be their partnership with 1Password and other password managers, Firefox, their partnership with the FBI, UK & Australian governments, etc.
What happens if Troy gets hit by a bus tomorrow?
And I think that's an extremely silly way to determine if something is a hobby project or not.
Also, just FYI, “The controller of the domain your email address is on will still see you in domain searches.”
Anyone wanting to do something nefarious with the emails will surely download the full source lists rather than try and scrape the aggregation site.
I wonder if he actually deletes the data...
Once a breach is determined, all of these passwords should be invalidated immediately and require a password reset if you're so behind you're not offering Passkeys or SSO. Rate limiting will slow credential spraying attacks, but the only way to eliminate them is to use SSO ("Login with") or Passkeys. You are negligent as a provider if you are not invalidating leaked credentials in a timely manner.
https://pages.nist.gov/800-63-FAQ/#q-b05
> “Verifiers SHOULD NOT require memorized secrets to be changed arbitrarily (e.g., periodically). However, verifiers SHALL force a change if there is evidence of compromise of the authenticator.”
> Users tend to choose weaker memorized secrets when they know that they will have to change them in the near future. When those changes do occur, they often select a secret that is similar to their old memorized secret by applying a set of common transformations such as increasing a number in the password. This practice provides a false sense of security if any of the previous secrets has been compromised since attackers can apply these same common transformations. But if there is evidence that the memorized secret has been compromised, such as by a breach of the verifier’s hashed password database or observed fraudulent activity, subscribers should be required to change their memorized secrets. However, this event-based change should occur rarely, so that they are less motivated to choose a weak secret with the knowledge that it will only be used for a limited period of time.
Does it really matter? I think all of my accounts use 20char autogenerated passwords from google that are unique for each account. So if one is breached, it’s just breached. Seems to have the same protection as a passkey.
Long strings in password managers was a shim until Passkeys got here, because passwords suck. This is a well worn path in enterprise with PKI. Passkeys are PKI for the Average Joe. Folks here will always have esoteric auth use cases, but you design for the average on this topic (consumer auth).
https://passkeys.2fa.directory/us/
https://bitwarden.com/blog/a-closer-look-at-password-statist...
> 19% of respondents said they used “password” as their password (!!!)
> 52% use easily identifiable information in their passwords, such as company/brand names, well-known song lyrics, pet names, and names of loved ones
> Best practices are still diluted by bad habits, with 85% reusing passwords across multiple sites and 58% relying on memory for their passwords
> A majority (68%) of respondents manage passwords for 10+ sites or apps and yet 84% of respondents reuse passwords
> More than half of respondents forget and reset their passwords on a regular basis
> Around a quarter (20%) were affected by breaches and a majority (80%) were prompted to reset their passwords
> Over half (56%) are excited about passwordless options, and 50% are using or would use ‘something you are’ forms of passwordless authentication
If you have concerns about Big Tech treating Passkeys in an anti competitive fashion, I would strongly encourage you to file a complaint with the FTC when that evidence is observed (as I mention in another comment here [1]). We need these primitives to deliver a better digital experience but also need to defend against fuckery using legal and regulatory mechanisms.
Amazon: https://www.aboutamazon.com/news/retail/amazon-passwordless-...
Uber: https://help.uber.com/riders/article/using-passkeys-to-sign-...
Ebay: https://www.ebay.com/help/account/signing-account/signing-ac...
Github: https://github.blog/2023-09-21-passkeys-are-generally-availa...
Link by Stripe: https://app.link.com/
Docusign: https://www.docusign.com/blog/docusign-customers-can-upgrade...
Tiktok: https://newsroom.tiktok.com/en-us/passkeys-fido-alliance (TikTok has over 1.677 billion users globally, out of which 1.1 billion are its monthly active users)
Google's Titan key now supports Passkeys if you need a secure hardware authenticator: https://www.wired.com/story/google-titan-security-key-passke... | https://store.google.com/us/product/titan_security_key?hl=en...
I like passkeys, they’re nice.
But I think you want to compare passkey users to complex password users.
I know lots of “normies” and they all just accept whatever their iPhone does. Which is creates a high entropy unique password for each site.
Yes, yes, I know, the site maintains that it's the victim's responsibility, prior to any bad actor taking advantage of the service, to sign up and then disable their information from showing up in the search.
Because the shock and awe of other users seeing 'Your info is out there!' immediately after entering their address, instead of after some kind of email verification, is more important than user safety.
If you’ve got someone’s email address, do you really need HIBP to figure out where it’s used? Wouldn’t Google give you a lot of results already?
At that point you just go grab the data breach from some other location and pull the data. Including passwords, which will often be reused across other services and give you access to other accounts. The breach data (or subsequent accounts accessed from the passwords) may have a lot of additional data, like location, personal preferences, message contents, etc..
It's _much_ more impactful than simply Googling someone's email address.
Do you have examples you can point to where this has happened? I would be interested in cases where this kind of approach has been taken, as my intuition is that Troy’s service wouldn’t make this significantly easier or cheaper.
Imagine you want data on a specific individual. You can locate and obtain ALL available data breaches, then search through each one individually looking for references to your target. OR, you can enter their email into a search and get back a list of which SPECIFIC data breaches have information on your target AND the types of information included in the breach. This narrows the scope of effort by orders of magnitude.
How does that not make it significantly easier and cheaper?
This enables bad actors to quickly and easily figure out which leaks they need to go look at to find further data related to the email address they searched on.
It'd be a great service if it hadn't spent the last decade deliberately ignoring a fundamental principle of security: Assume there will be bad actors.
It takes a bit of cognitive dissonance to leave the problem unaddressed for so long. eg, if you assume there are no bad actors, there's no need for the service to exist. And if you assume bad actors DO exist, the service needs to protect its core functionality better against being an engine for abuse.
Is Troy Hunt a Mobius variant?
All my stuff's been locked down with a password manager/2fa so I'm not worried, but having been on the internet for ages it's pretty funny at this point.
Adobe.com
Bit.ly
Bytargentina.com
Cafepress.com
Chegg.com
Contentful.com (via Apollo)
Dailymotion.com
Disqus.com
Dropbox.com
Edmodo.com
Facebook.com (via Zynga)
Gawker.com
Invisionapp.com (via Apollo)
Kickstarter.com
Last.fm
Linkedin.com
Linux-mag.com (via QuinStreet)
LiveAuctioneers.com
Monster.com (via Apollo)
Myfitnesspal.com
Parkmobile.com
Streeteasy.com
Teespring.com
Ticketfly.com
Ticketmaster.com (via Ticketfly)
Tumblr.com
Xbmc.com (via Kodi)The last time this was asked on HN (Feb 2018) I had 13 breaches and appeared in one “paste” - https://news.ycombinator.com/item?id=16465030
This returns a list of possible suffixes which can be checked for the actual password to see how many have been breached.
For example, a search for "abc" with the hash "a9993e3...89d" becomes:
`curl -s https://api.pwnedpasswords.com/range/A9993 | grep -i e364706816aba3e25717850c26c9cd0d89d`
which returns `E364706816ABA3E25717850C26C9CD0D89D:226273` indicating that the password has been seen 226,273 times
My e-mail address (after requiring JS and cookies) was allegedly found in an collection of avatars ...
I'm sure it's now in a collection I rather not have it in.