Why It Matters Whether Hashed Passwords Are Personal Information Under U.S. Law
jdsupra.com
jdsupra.com
But salts aren't used to be a secret: they are used to make existing rainbow tables useless, as the salt wasn't used to precompute the rainbow tables.
In the example they use "2@3" pre- and post-appended to the password. Ok, so any rainbow table that haven't used 2@3 is useless. But "password" will always be a lousy password and 6 characters of salt isn't that great.
I should say correct enough.
(This is still awful, because then two users with the same password would have the same digest. If you know that joe@example.com has password "badpass", and jane@example.com has the same hashed password, then hers is also "badpass". So use randomly per-user salts, gang! Even better, pick a tested implementation of PBKDF2 or Argon2 and use that implementation. Don't roll your own crypto!)
But still, it's better than nothing.
Just pointing out that I have seen conditions where the article is not terribly wrong in its statement.
Things like bcrypt are nice because they take care of that for you.
It's the same logic as applying well known target data like birthdays, anniversaries, etc.
I was talking about _reversed_ passwords, guesses that fulfill the algorithm. Bonus points if they look like names, words, etc, but even random junk might just be a pwgen password.
Discussion from 2017: https://hashcat.net/forum/thread-6219.html
From 2018: https://security.stackexchange.com/questions/186953/cracking...
GPUs are much faster now. I doubt you’ll find many people using rainbowcrack or similar tools in 2021.
Wouldn't it be a pepper then?
https://en.wikipedia.org/wiki/Pepper_(cryptography)
edit: nvm, it seems the main difference is a pepper being secret.
Salts must be unique per password as far as I know? Otherwise they don't do their main goal which is to make every guess useful against only one hash. (ie: guess 'hunter2', check if any user had hunter2 as a password, but salted, you'd only be able to guess if cschneid had hunter2 since you'd be hashing hunter2-sssssaaaaallllltttt).
If right or wrong is a spectrum.
If the attacker hasn't seen the code, they won't know whether the salt is prepended, or appended, or both to the password (or if there are two salts), but the amount of computation needed to try all possible salts is roughly equivalent to the amount needed to crack a single well-chosen (unsalted) password.
PBKDF2, obviously a less marketable name than scrypt, was a pretty standard tool for dealing with fast hash functions. What turned out to be a bigger problem than speed is memory use. Now that we have widespread GPUs, it's a lot easier to parallelize password cracking, and most cryptographic hash functions use negligible memory, so it parallelizes very well. scrypt also improves this.
I actually saw scrypt take down a website. It got attacked with brute-force password attempts (<100k, so it wasn't a serious attack), and the scrypt calls exhausted the CPU, essentially doing a DoS attack. We ended up wrapping the call in a semaphore to prevent this. You don't normally think about it, but login might be the most expensive operation your webserver does.
At the very least use a well-studied iterated salted hash algorithm for storing passwords (if you must authorize an incoming request later). That is a minimum bar. I generally recommend argon2. Reasonable alternatives are bcrypt, scrypt, and PBKDF2. PBKDF2 is weak against hardware, but it's still better than using a hash directly.
I teach a class on developing secure software, and students have to do a final project. I take off points if they directly use SHA for password storage; that's just not acceptable.
Cryptographic hash functions are typically designed to be computed quickly, so it is possible to try guessed passwords at high rates. You could try billions possible passwords each second. This is why we have password hash functions that perform key stretching, such as PBKDF2, scrypt, or Argon2. They increase the time (and sometimes even memory) required to perform brute-force attacks on stored password hash digests. Use a large (256-bit is fine) random salt value. All the salts are random values (note: they do not have to be a secret), so each user will use a different salt value, and now the attacker has to compute the stretching function once for each password combination, rather than once for each password, and this is a lot more work for the attacker.
If you only hash with a salt, it still leaves passwords exposed to brute-force attacks and dictionary attacks that they can easily run on GPUs. Key stretching is important! This is why we have dedicated password hashing functions that are slow enough to mitigate brute-force and dictionary attacks.
TL;DR: salting and stretching on passwords, along with password hash functions (slow password hashing functions: PBKDF2, bcrypt, scrypt, Argon2, Balloon and some recent modes of Unix crypt)! Do not use fast cryptographic hash functions on passwords, as it defeats the purpose of stretching. Salting is not enough.
The correct primitive to use is a key derivation function. Examples are PBKDF2, bcrypt, scrypt and Argon2.
A good starting point when trying to decide what crypto algorithm to use is https://latacora.micro.blog/2018/04/03/cryptographic-right-a...
Your standard crypto library probably implements it (Rfc2898DeriveBytes for .NET). Under the covers you should tell it to use SHA512. If you can choose something better go ahead, but this is often good enough
If people use weak passwords, the hash algorithm doesn't matter, they're going to get cracked. If you use strong passwords, the hash algorithm doesn't matter, they're not going to get cracked. This sort of thing only matters for medium passwords. And can be avoided by requiring strong passwords.
It also does very little against precomputation attacks (e.g. rainbow tables), because it makes it take longer to compute the tables but it's still efficient to use them.
Meanwhile it adds a user-perceptible amount of latency to logging in, because the hash algorithm is so slow. That causes users to take steps to avoid it, e.g. by staying logged in when they don't need to or preventing their machine from locking, which reduces security.
At that point it may be suitable to store the password as plaintext, though a layer of hashing could be suitable to slow down an attack from authenticating in case they get read access to the database
It also doesn't have to be 64 characters. 16 characters in base64 is 80 bits of entropy. That's hundreds of years worth of hashes on a Bitcoin ASIC drawing hundreds of watts, for one password.
If these distinctions are new to people on this thread, I recommend https://www.oreilly.com/library/view/building-an-anonymizati... for some background.
Even though I generate a random salt when hashing passwords I wouldn't use it to select a user out of a database because I don't check for collisions by ensuring a unique salt that has never been used before. The chance of a collision is tiny, but I still don't want to do it. But this would still qualify as unique under data regulations?
The policy view of what these things mean vs. what we talk about in terms of entropy, collisions, confusion, etc, are subtly different things. Of course it depends on the regs regime, but here's an example from what I do:
A business wants to share 10million of personal profiles keyed on a SIN with another business or agency for some research. Privacy law says they can't share the data unless the personally identifying information is removed. Business tells their DBA to "encrypt" the SIN numbers, who says "sure" and SHA256's them all and shares the data set, because to him, now the SINs are not shared(!).
A privacy policy analyst freaks out because hashing the SINs has done nothing to protect the identities of the people in the data profiles. The DBA can't figure out why because he used an NIST approved 256bit hash on them and he tells his boss "it's fine, they're encrypted."
To your case with the salt: if you salt the hashes, the agency receiving the file calls back and says they can't use the data because the hashes don't match the profiles in their database - because they have been using the hashes as unique identifiers and re-identifying people in the data set. If you salt them for your initial data sharing, and then re-hash them with a new salt for an update, that's better, but it is still a transferable identifier for that cut of the data.
Even if you tokenize each profile with a UUID and transfer that UUID between agencies, you are transferring a unique identifier about the person. The right way to do this is to have a tokenization service broker that takes records and synthesizes new record keys (MBUN) for each destination counterparty you are sharing the data set with. Hardly anyone does this, and they just take the privacy risk instead.
The policy vs. info theory difference is that policy can conflate hashing and encryption depending on the purpose of the use, where in tech and security, they are totally different things.
In the practical sense what matters is whether you can safely assume the hash for a particular user in a particular system will ever be unique.
E.g., if I had a snapshot of a database mapping identities to password hashes and a log of all the hashes that had been computed by the auth server, I could make very reasonable guesses about who logged in at what time.
With that said, I am not a lawyer and I have no idea what the legal significance of all that might be.
Your phrasing was: "If the hashed password *is used as* a lookup key, that makes it a unique identifier, and it's PII."
If I take your wording literally, your saying that I could store everyone's SSN in a database, but as long as my system didn't attempt to use SSN as a unique identifier, the SSN doesn't count as PII.
But it seems unlikely that that's how you meant it.
Privacy is anti-discrimination, and the reason social media companies are so rich is because they sell micro-discrimination as a service. It's so valuable because when people see how it works they ask, "how is this even legal?"
I want to put a vote in here for an anchoring. Lots of methods:
1) Code secret. This is ok but loses all strength if disclosed.
2) KMS wrapping of hashes. This is better; attacker needs to be in your environment, actively unwrapping all your stuff.
3) Controlled hashing step (best). Have one step of your password hashing incorporate a one-way operation using a "normally non-extractable" hardware secret. This could be in CloudHSM if you're in AWS.
This completely prevents offline attacks unless the attacker compromises your HSM.
4) It would be nice if cloud providers implemented "password hashing as a service". They could easily and cheaply prevent offline attacks.
5) Of course strong key auth instead of passwords for clients would be nicer.
1. This article repeats the misunderstanding that Rainbow Tables are the same thing as a Dictionary:
> Given the processing power available, hackers have and continue to generate enormous tables that contain anything from every possible combination of values for shorter passwords to lists including variants of common and known passwords. These are called rainbow tables, and if the password used is included in one of these tables, then the cleartext password is known to the hacker.
The Rainbow Table is actually an interesting improvement on a pre-existing idea (by Hellman, yes as in Diffie-Hellman) for a Time-space trade off attack.
So let's describe the most basic idea: We have a list of potential passwords (a "Dictionary"). It doesn't matter whether it's an actual written list, or a procedure to pick possible passwords, but it must be strictly finite. Just trying them all is a Dictionary attack. For an online service this is one thing "rate limiting" is defending against.
Now, if you have the hashes, you could do your dictionary attack except instead of trying to log in you just calculate each hash and compare it. There are automated tools that can do this for you e.g. John The Ripper.
Next cleverest idea is: Let's calculate hashes from our dictionary, store and index them (or we could e.g. store them in order) with the original word so we can just look up a hash and get back the password if we know one.
Hellman's idea was: Don't store all these hashes. Storage is expensive. Imagine a function that converts a hash back into a password from the dictionary, call it F() and the hash function H() now instead of calculating and storing H(password) -> password we recurse through H(F(H(F(.... some number of times, effectively chaining together many solutions in each entry stored, then we do more work to calculate each intermediate when doing a lookup of a hash later. We're trading off less space for more time.
Hellman's idea is clever, but it's a bit wasteful, because sometimes there will be (at random) collisions in either H() or F() and this means you're wasting space/ losing recall in your system.
So, Rainbow Tables are Philippe Oechslin's improvement, Oechslin deliberately perturbs F() differently for each iteration so that if a collision does occur it will go away again unless it was in the same iteration of the chains.
This produces a "rainbow" picture in Oechslin's slides which named the resulting technique.
Rainbow Tables are useless against a properly salted hash, because all this up-front work must be done again for each different salt, or, the size of the tables are inflated by the size of the salt. So e.g. a 32-bit salt (many today are far larger) makes your Rainbow Table cost four billion times more to calculate and need four billion times as much space.
They got famous because some key Microsoft password technologies from the turn of the century do not use salt, and so you can attack those with Rainbow Tables. In particular the LM Hash can 100% reliably be reversed by a suitable Rainbow Table that would be affordable to store in that era.
2. It gets salt wrong:
> As long as the salt value remains secret, this is a very effective method against rainbow attacks and other current methods of attack.
Nope. The salt doesn't need to be secret it should be random for each record. Our goal isn't that an attacker doesn't know each salt, which in most of these cases will have been stored with the hashed passwords, but to prevent them amortizing work over a large number of hashes, see the discussion of Rainbow Tables earlier.
Overall I think for lawyers or managers looking for legal advice, which I assume is the target of this article, the focus ought to be on getting away from passwords. They were already a bad idea, a last resort, last century and we've doubled down rather than seeking every other alternative.
If your system is authenticated with WebAuthn you do not store any secrets you do not store anything that would "permit access" to the system, because it's all driven by Public Key Cryptography. You store things can be used to verify that the client is the same as before, but they can't be used to pretend to be that client, or even to unmask the client's identity. You could (but probably shouldn't) paint the contents of a WebAuthn authentication database on the side of your building and that'd be fine.
Unfortunately "good" passwords have other problems, like being hard to type and remember... so there are practically no good passwords out there...
Isn't this already answered by existing law? Wasn't Kevin Mitnick already charged and prosecuted with laws that would cover unauthorized access? Obviously if you have a hashed password and it's not your password and you're trying to access an account that isn't yours then I'm sure that qualifies as unauthorized access. What am I missing here?
It sounds like there is an argument being put forth to make hashed/salted passes PII for, I can only guess, the sole purpose of litigation against businesses for being hacked. Otherwise, does anyone think hackers care about PII?
I can understand prosecuting for security lapses that are related to a data breach. But within reason. You're getting close to charging the victim for the crime here.
I know it's fashionable to argue against storing any PII here on HN. But that's not reasonable for our society. Because, much like salting+hashing, when you're implementing security measure X, all the future hacker has to do is X+1 to breach the wall. This game will never end. If we go down this route, you know what's going to happen? Only Amazon and Google and Facebook can control your PII. With CCPA and GDPR you're already seeing how this plays out. The businesses with deep pockets win. They can afford to deal with the legal realities of doing business all day long. There is an entire cottage industry that popped up just to put those stupid cookie banners on websites now. Does anyone really want to live in a world where Google gets to dictate if your business can exist, and Amazon constantly stepping on your neck and demanding rent?
CCPA/GDPR are basically laws that punish the entire world for the sins of three or four monopolies.