In an offline brute forcing scenario, either the passwords have been hashed with a strong key derivation function and a randomized salt or they haven't. If they have, it doesn't matter if the password is in that list. If they haven't, you're most likely screwed either way, because attackers can get up to 1 trillion password attempts per second in real world cracking setups now.
Extrapolating further, this "policy" would disallow people from choosing passwords below a certain length, because it's theoretically plausible someone has published the set of all possible strings fewer than n characters long.
Yes, that seems like a good password policy. A list of possible alphanumeric strings that is actually reasonable to physically publish (i.e., not 20 trillion) is a list of extremely short alphanumeric strings. 5 alphanumeric characters is 380 million possible strings. 6 is about 2 billion. You should absolutely ban passwords that are 6 characters or shorter!
In fact, I would go so far as to say that the questions of "Is this password too short because someone could brute-force all the possibilities, even if we're using a good password hash and a previously-unknown salt" and "Can someone physically enumerate all passwords of this size and put them on Pastebin or otherwise get them in the HIBP database" are equivalent.
And this is aside from the fact that you won't even rip through those 20 trillion in an online brute force attempt. Eventually you are only constrained by computational resources. If the passwords are cryptographically secured, it's fine if any one of them is published on the internet if you cannot associate it with any given user. If they're not cryptographically secured, this won't meaningfully reduce your already poor security anyway.
My principle is that you should not let people sign up with breached passwords at all - don't make judgment calls about which breaches matter, and whether you think it's the same user or not, or the password is strong enough or not. Just ban the passwords. (Remember that no actual data breach contains 20 trillion passwords.)
The world has decided that making me go through a Google CAPTCHA when registering a new account - a process that takes me several seconds of active mental effort - is fine. If the check takes even 1 second of server time, is that noticeable?
There are MUCH smarter ways to implement this than grepping a list line by line. I mean seriously who the hell would do that... If your bio isn't complete BS then surely you understand that proposed implementation is BS.
That being said I'm actually going to concede this argument, because on further investigation Troy Hunt provides the entire copy of his database freely for local lookups. Once you hash the user password to match the database hashes you can introduce the optimizations you mentioned, and in any case it shouldn't introduce intolerable latency to do that lookup locally.
Blocking a list of 20 trillion passwords is probably overkill if you have a slow hash. But with a fast hash it's the difference between "impossible" and "less than one GPU-week".
And if it's easy to block, you might as well do it. There's no upside to letting people use already-posted 20 character strings.
If you publish one trillion passwords from a large space then each one of them gets a probability boost (of approximately 1/1 trillion), though not enough to ban them, especially if they are not actually associated with accounts.
The danger with using a rare but breached password is that there is actually quite a high chance that it was breached from your account elsewhere.
I do agree that if the entire space of possible passwords is only 20 trillion, that doesn't change the probabilities. But there are over 20 trillion eight-character alphanumeric passwords. So, I would actually say you should ban them all, because you should insist your passwords are at least eight characters long. :-)
Edit: yes, agree, in practice the probability boost is not very much. I'm just saying you may as well ban them on the assumption that the HIBP API will do so at its current level of performance. (20 trillion is a ridiculous number, because it's much larger than any possible breach and yet much smaller than any meaningful password space, so any arguments about it are going to be inherently silly in some fashion. My current silly assumption is that the HIBP API is capable of ingesting 20 trillion breached passwords with no performance hit.)
If not then it really doesn't matter that they got published. They're useless to hackers without knowing what email to type in. The attack model is that the attacker actually has to log into a website and you don't get 20 trillion attempts.
> The attack model is that the attacker actually has to log into a website
Not to find the password. If it was then nobody would get upset about plaintext password storage.
But anyway if you pick a password from a list of 20 trillion where the offline attacker knows the list, it doesn't actually help them much because a single selection from 20 trillion options has 44 bits of entropy.
Passwords that users choose typically have less entropy than that afaik
No because they each still have a very low probability.
You have to be Bayesian about this: a list of one trillion passwords that have no further distinguishing information about each one of them cannot be assigned a probability of > 1/(1 trillion)
In a data breach, a given password appears next to a particular username or email, which means it has a very high probability of being the password for that account.
Remember the "Debian weak keys"? What's weak about those particular keys?
Nothing. Nothing whatsoever. Those keys aren't special in any way. Except, Debian shipped releases that always picked one of these key pairs. So anyone with a mind to can go find the list of private keys that corresponds to these particular public keys and thus we don't let you use those public keys any more.
(You can go try this if you don't believe me, submit a CSR to your preferred public CA asking for a certificate for one of the Debian Weak Keys, it will be rejected and there may or may not be an explanation attached saying your keys are crap and to get new ones)
Whole swathes of keys are blacklisted. ROCA is another example, somebody took one mathematical short-cut too many in their optimised design for RSA key generation, and so the resulting keys all have this very obvious structure that's exploitable (not easily, but enough that a sovereign entity could definitely break them). So we just blacklisted all those keys.
If you pick truly random keys you'll never notice this in a lifetime because of statistics, it's just some code on the issuer's systems that you never need to care about.