I'm Open Sourcing the Have I Been Pwned Code Base
troyhunt.com
troyhunt.com
> HIBP isn't in a state to simply flick the visibility of it in GitHub, but it needs to get to that point. Instead, I need to choose the right parts of the project to open up in the right way at the right time.
and then further down:
> I don't have a timeline for each step along the way yet as HIBP remains something I do in my spare time and I've always got a bunch of other stuff on my plate, but the process has already begun and I'll be sharing more on that as soon as I can.
Intel's 20GB data leak was on HN yesterday, with a hint that more is coming. Will be interesting to watch ;-)
In Signal's case, they could just let you send messages that might never be received because the contact doesn't currently have the app. Also gives the receiver plausible deniability unless they respond. And could help get more people on Signal if they did the "Someone sent you a message, download the app now" trick.
Of course, if you can’t verify what’s running in your server then neither can your users. All you can do is be as transparent as possible about how you process their data and hope that’s enough to earn their trust.
As an aside, homomorphic encryption has been touted for a while now as a way for services to process encrypted user data without ever being able to decrypt it. Haven’t seen much use of it in the industry though, last I heard it was enormously impractical.
On the other hand, users need to put their trust in humans, not computers.
Arguably, the probability of such an attack is extremely low--slim to none might be an overstatement--but it's an interesting thought exercise.
But have can we know the "trusted compiler" is really to be trusted?
On the other hand, users need to put their trust in humans, not computers.
It’s still a human trust problem all the way down because you had to buy your computer from someone, and they had to buy the parts from someone, and so on.
There are some places where it’s less of an issue due to cryptographic magic, but those are few and far between.
So here it's taken one step further, you don't have to trust that it's the same code running on HIBP, because you don't have to trust them at all for the service to work.
It’s possible some bad actor could do dumb stuff, but SSL helps in that at least the server would have a unique id.
Personally the first time I used HIBP I checked the javascript that was running and read about what happened.
It’s possible that someone could hack their server and replace the javascript with new bad versions. I don’t check anymore. But it’s unlikely, and a shared, common risk of using any site, I think.
But the important thing with HIBP (and other sites) is that I never send them any sensitive information.
I couldn't agree more. Even if you have root on their server and do a full audit, there's nothing to stop them from changing the code the moment you log out.
That's a valid criticism. It only takes one person finding clear evidence of problematic behavior to advertise that fact to the entire community. So long as a small fraction of people do actually check, the whole community will be fine. But if only a negligible number of people check it, then perhaps no one will be checking when it is abused.
> And again, how do you prevent that different users get different functionality?
Well, it depends on how the system is discriminating among users. To a great extent, that kind of abuse is prevented by anonymous users. You don't have to log into HIBP in order to use it. The system could still be discriminating by IP address, or even by various kinds of browser fingerprinting (including just collecting the advertising company cookies that so effectively eliminate anonymity for much of the web). Occasionally using TOR won't help against this -- a malicious operator could simply provide "clean" functionality to all TOR users.
- - - -
On the whole, things like, k-anonymous API design, making source code open, or having a small fraction of users checking for security issues all help a great deal and make it difficult for a malicious operator to abuse the system. But none of them are perfect. In the end it comes down to trust. Troy Hunt has EARNED that trust due to his openness, and in no small part by choosing to implement protections like these that he didn't need to implement, and I find that extremely persuasive. If you take steps to ensure that your CANNOT cheat your "customers", then that's pretty good evidence you aren't likely to be trying to cheat them.
I investigated this thoroughly a few years ago, and the solution I came to is distributed computing. Join a peer-to-peer network an run the code directly on your own hardware or depend on someone that you actually trust. As far as I can tell there are no shortcuts that allow you to trust strangers with sensitive data.
If your password is in the list, you should probably be changing. It being bad is how it ended up in the list.
According to [1] by Troy Hunt, around 86% of passwords from one dump were already in the HIBP database, so attackers could assume your password is in the list with 86% certainty, extrapolating from that.
If we look at the returned data for a sample hash prefix (that of "password", which is 5BAA6), and sort them by count of password use, we get these password hashes with the most usages:
~ curl -s 'https://api.pwnedpasswords.com/range/5BAA6' | sort -t: -k2 -n | tail
42BAADCD710F9EA7E62B60E01D05469AC64:14
EA2008F79BE2B0E0C02A1642725433BBB2F:15
3A8ADE4CF1DAD5342AF2F9FC9247EC21943:18
5E2BCB2FEF09257B0306B4744418999611B:18
A516C42C8CD4C7E7E328ABB90D002A9890E:29
CF2F87E596758D031C0006D1827C9908E5C:34
EF0E14CCB17E525D76050283148A57828F8:44
8E0D5C9D144BACC76E52C44F5B61E8DF629:213
2648FB0B2EDA4FDFF99BF51E912CD95C023:7201
1E4C9B93F3F0682250B6CF8331B7EE68FD8:3759315
Summing the counts gives 3768295, so 3759315/3768295= about 99.7% of passwords with that hash are "password".Going with a more obscure password, like say "obscure", you get 1801/4319 = about 42% of passwords being "obscure". This means that if I search for the password "obscure" on HIBP, HIBP can be about 40% sure that my password is "obscure".
How is this not a huge security issue? I would agree that folks with unique passwords not appearing the database are safe, but anyone using a common password can be identified with pretty high probability, and even someone with a password only occurring once in the database is at risk because a typical query returns about 500 results: low enough that a human could brute force a web input that didn't have rate limiting in an hour or two.
[1]: https://www.troyhunt.com/86-of-passwords-are-terrible-and-ot...
If you personally want to check your password wouldn't it help to curl more queries? Probably also taking care to make not too much per some unit of time?
I also see problematic that, if I understood correctly, some browsers query automatically that service with your passwords which browser collected whenever you “saved” some login.
And they claim that is “for your security.”
If you have a match, you will certanly use different password.
And, by the way, they have username/hashes before you even use their API.
I agree - this is the main use of HIBP.
> If you have a match, you will certanly use different password.
I doubt this is true in all cases. Sometimes people get lazy. But nonetheless, a malicious actor could still capitalize on the few minutes between a user checking a password on HIBP and that user changing their password. For example: for every password lookup, pick the most likely password with that hash, and if it has at least 90% share of the passwords with that hash, try to correlate an email address to their IP address (using leaked dumps), and then try to log into gmail with those credentials. Even if it works 1% of the time, a lot of harm could be done in a few minutes. (For example, US Social Security Numbers may appear in a years-old employment document in the email).
> And, by the way, they have username/hashes before you even use their API.
But they don't necessarily have yours. If I am somebody who uses a different password on every site, and I check a password for a site that has not had a leak, then they do not have an association between my email and my password before my query, and they do after my query.
Regardless of the particulars of possible exploits, I'm mainly claiming that the original comment above (by @matsemann) is not true, in that some level of trust in HIBP is necessary to input a sensitive password.
Ok so many use cases are removed - maybe your locally running password manager allows you to test a potential new password against leaks. Now we have a rare example of what you're hypothesizing, at this point you're trying to tie an IPv4 address to an identity. Many networks have either shared IPv4 addresses (NAT/GNAT in major metros, corporations, public Wifi), or dynamic IPv4 addresses (change daily/weekly/monthly). It's pretty hard (though not impossible) to link an IP to a person.. and vice versa, between 3g, 4g, 5g, wifi, work, home, complementary, wired and wireless connections a single person may appear across several addresses.
Finally the idea of accounts - knowing a specific IP just tested the password `password` and Aretha Franklin is most likely at that IP.. well, which of the hundreds of services might she be potentially considering the new password for? (again - testing, not setting the password to). If you could narrow it to one service, or brute force all of them, and assume the user ignores the dire warnings, you still need to know their user credential (be it email - of which they may have many, username - of which they may have many, or service-generated username which you'd need insider knowledge to obtain)
If you're worried about trust of your sensitive password, your password manager and the service you're using it on (if you reuse passwords) need far more.
Do not run any of your own application logic on the server - your javascript client should talk to the standardized datastore directly.
Then your users can see all the logic on their side, and inspect all the data, and if they want they can host the client side stuff themselves and run their own server.
In this model, the client must validate the data before it puts it into the data store, and again when it gets it out of the data store (since the data could have been put there by a malicious/modified client)
On the other hand, if you are bothering to email me to ask for SSH access to ensure what I say is true, you probably have the knowledge to detect if I am lying from inside the server.
Troy got it; so can you.
Both as a curiosity, a research tool, and a crime tool.
As a hobbyist, I used to have a folder of breached data (ashley Madison was really fun). And if you add in groups like /r/datahoarders, the data will probably be kept someone.
The dataset could be reconstructed if needed.
They almost never tell you which site was actually breached, nor do they ever give you any hints as to what password was actually compromised.
So really, when I show up in one of these alerts, I'm always asking myself:
- Was this a recent breach, or a redundant alert from something I dealt with months ago?
- What account do I actually need to update, if any, to be safe from this alert?
These questions almost never seem to be answered. As someone who uses a password manager and a different random password for every site, there's no way I'm going to proactively hunt down and change every single entry in its DB when I get the "alert of the week."
(FWIW, I once worked somewhere that somehow had access to a far better version of this data than they'll ever let the public get access to. That system actually did generate alerts that were actionable. I only wish I had a way to get useful alerts like that as a private individual.)
If you really want to hunt it down, a lot of this data is available via torrent, although probably not too useful for generating alerts.
[1] keepassx does not give a way of searching passwords. You can manually look through the list of passwords though.
That's probably true for most people. But that's what email aliases are for. If I notice (through 1Password alerts) that my me+facebook@mydomain.com email got leaked, I pretty much get an idea which site was breached. At least GMail and Fastmail support email aliasing.
In grad school we looked at integrating CT monitors into web servers that manage certificates for your sites (including monitoring lookalike or spoofy names), but then weren't sure what to do when a suspicious certificate appeared in the logs. Do you email the site owner? Then what? Sure, you can report to CAs and web hosts and all that, but who will actually go to that trouble? By the time you do that, the site will probably already be blocked by SafeBrowsing and whatever other blocklists.
Last time I used HIBP, it told me what site the leak was from.
If your password shows up at all, you should change it wherever you use it.
48 hours before its on HN.
You just have an open VM
For something like this, keeping things private has allowed Troy to work out the best way to do things with complete autonomy. That's really useful if you have the time and resources to get to market on your own.
Now it's stable and best security practices are nailed down, he can open it up to a bit of scrutiny and feature bloom.
You'd do it to prevent feedback.
Think something like he started the project at some point, never planning for it to become this big. It was probably some code written in a random directory, with no regard for whether test fragments are around, whether it contains test data that maybe shouldn't be public, different pieces that are closely tied to a server config that itself is very specific to what he already had running and would need some documentation to be useful etc. pp.
Isn't the code just essentially a text input box that takes a string, hashes it, and runs it against hashed passwords in a database?
He's open-sourcing it so the community can help him manage the project, which I think is a fair request/hope.
He touches on the non technical difficulties as well with his comment: "We invite parties to form their own views on the legality of the data." So the fact he's gathered it all lets HIBP be an service that other companies can use without worrying about the thorny legal question.
But there's also his reputation as a steward of the system, which is valuable beyond the data itself.
Anyway, while he didn't actually open source anything yet, I'm glad he's committing to it, as hopefully that will allow this internet resource to continue.
"and it took a failed M&A process to get here"
Second paragraph:
"especially in the wake of the M&A process[0] that ended earlier this year right back where I'd started"
[0]: https://www.troyhunt.com/project-svalbard-have-i-been-pwned-...
https://threatpost.com/troy-hunt-sell-have-i-been-pwnd/14556...
Maybe has efficient hash comparisons or something...
https://threatpost.com/have-i-been-pwned-no-longer-for-sale/...
Socialism for "other companies," capitalism for us. It's details like this that prove to me how purely ideological it is to claim that we need capitalism to "produce value." _We_ produce value already. Capitalism uses that value for free or a ridiculously low rate, then turns around and charges us for it.