Storing Your Data Securely: A Primer
tinfoilsecurity.com
tinfoilsecurity.com
Having my email address on my blog is a vulnerability. Another vulnerability is named 'Rudimentary Scan', meaning 'we looked, and your site has no client-side vulnerabilities'. These issues mean my blog is 'borderline unsafe'. Wut.
Neither of these are vulnerabilities; to call them such does a disservice to your customer's understanding of technical language.
Rudimentary Scan means you ran a scan from our homepage, pre-verification, and we couldn't run our full suite of modules. All that is saying is that now since you're verified you should re-run a scan which will look for quite a bit more. We figured putting it there would be the best way of not giving a false sense of security from our 'rudimentary scan.' We make it 'borderline unsafe' so as to prevent that false sense of security. I can totally understand your confusion, and we should make this clearer. Ideas / feedback are always appreciated. :)
The email address is an informational vulnerability, which is dismissible.
Does that clarify things? (cofounder of Tinfoil)
Say I encrypt the entries... Since the database runs in local host, if the host is compromised and subsequently 'rooted', what exactly does database encryption offer?
The intruder will sniff/find the encryption key(s) since the key must be somewhere inside the application or in memory (don't know how, but I'm sure it's possible) in order to be able to decrypt data on the fly. The way I see it it's just added complexity for the admin with no gain.
Is there any way of defending an SQL database even when the host is compromised?
Separating concerns be user account has its uses (e.g. sandboxing an app from accidentally screwing with files it doesn't need access to), but I don't think you necessarily want it as the linchpin of your encryption strategy.
> but I don't think you necessarily want it as the linchpin of your encryption strategy.
Its not about having one thing at "the linchpin". In fact, you should not rely on any singular thing, but rather separate out pieces of information so that no part has more than it needs. The test user should not have access to the production database and vice versa. In sum, you should use multiple things to defend yourself, not just one.
The best answer I have to this problem is a variant of the HSM: the server stores ciphertext but not the key. To decrypt data, it's presented to another server, whose only purpose is to "seal" and "unseal" the data. That server:
* Doesn't run on shared hardware
* Has a totally different management process from the rest of the app servers
* Speaks only to authenticated, trusted hosts
* Is carefully monitored for usage anomalies
This isn't a panacea; in fact, its security value is surprisingly limited. But it does solve some actual problems.
KDFs are not for storing password hashes, they are for key derivation. This is a subtle point, the main thing is to use KDFs as KDFs (to derive keys that are then used to encrypt data). Details in the first link here (the second is about common mistakes when using scrypt):
- http://blog.ircmaxell.com/2014/03/why-i-dont-recommend-scryp... (bad title, he's referring to pass storage, not KDF usage)
- http://vnhacker.blogspot.com/2014/04/fairy-tales-in-password...
Another consideration is plausible deniability (PD) for situations where you're compelled to disclose your password(s). Our company writes Mac encryption software that specializes in this (and it uses scrypt). Here's a list of tools that offer PD (ours is called Espionage):
https://en.wikipedia.org/wiki/Deniable_encryption#Software
I compared Espionage's PD to TrueCrypt's on reddit:
http://www.reddit.com/r/security/comments/2b5icu/major_advan...
The second criticizes an API design with the scrypt code.
Neither makes a case that password-based KDFs are "not for storing password hashes".
Using a good password KDF as a password authenticator is a fine decision.
Yes, I didn't claim otherwise.
> Neither makes a case that password-based KDFs are "not for storing password hashes".
The first one did quite explicitly for scrypt: "And that's why I don't recommend it for password storage."
> Using a good password KDF as a password authenticator is a fine decision.
Apparently not so for scrypt, and scrypt is the best KDF I'm aware of (heard something about yescrypt being better but haven't had a chance to look into it).
I don't see why a KDF is needed for password storage. What's wrong with SHA256 + random salt? If for some reason you want a slow hash function, use the salt MOD some constant to do several rounds of SHA256 (though I don't see why that would be necessary).
You're referring to simple dictionary attacks on a large database of passwords? Even bcrypt won't help much with a password from the top 10k list. I wouldn't say that stopping rainbow tables is "virtually no protection". Please explain what you're referring to.
[edit]
Does anyone know if the bitcoin ASICS have been repurposed as SHA256 crackers? md5 and sha-1 are more common in password storage, but if you use what the parent suggests, then it could be worse since there is so much specialized SHA256 hardware out there.
You misunderstand me. Please have a look at my reply to pbsd, I think it addresses your comments there too:
https://news.ycombinator.com/item?id=8087975
I did say there: "I stand corrected about KDFs not being OK for password storage".
At this point, given the content on this thread, it seems that it's been explained clearly enough why salted SHAx is inadequate to that task. Do you still not understand why that is?
Until Windows finally got plugged into the Internet and the weak LANMAN hash became relevant to hackers, nobody used "rainbow tables" to break passwords (classic Unix crypt passwords aren't particularly vulnerable to them, since they're salted). But --- and this is crucial --- everyone cracked Unix password files. All of the original password cracking tools, from Crack through JtR, were designed to quickly crack salted hashes.
Don't worry, I understand very well (this is not rocket science :P).
My problem, I think, was an inappropriate attitude toward the realities of stored password hashes on servers, in addition to having read that "scrypt is bad for password storage" (and having misunderstood the "badness" of it). I had read too many articles that treated HASH + SALT as "OK" for servers that stored user passwords, and I hadn't kept fresh in my mind the speed with which these passwords can be cracked today.
For example, take the article I linked to in my reply to pbsd: https://crackstation.net/hashing-security.htm (recently written).
This article, and many many others like it, are written by seemingly reputable people, and yet their attitude is that HASH + SALT is OK for servers. Even though it discusses bcrypt, etc., it frames it as though it's an "optional" hardening technique, and places HASH + SALT under the category of "proper hashing":
https://crackstation.net/hashing-security.htm#properhashing
I haven't seen them called out on it. I agree with you and the others in this thread, that it isn't good advice. If 10-char passwords can be protected better with KDFs, then they should be.
I don't know what the "security community" is, but if it has any relation to "the community of people who can speak with any authority on cryptography", let me ruefully assure you that it is much, much smaller than you think it is.
Thanks for making that very clear. :)
Re "security community", sorry, I'd edited that part out before I saw your reply, and replaced it with a call for these posts to be called out for giving bad advice.
The main difference between a password-based KDF and a password hashing function is that the KDF usually needs to support variable-length output. Otherwise they perform the same task: take a secret string of low-entropy s, and stretch the cost of breaking it to 2^(s + w) for some work factor w. In other words, a KDF is a superset of a password hashing function, and it's OK to use it as one.
I mentioned I'd heard it was a good KDF (not a good password storage function). I didn't know that was its main purpose. Thanks, will need to read more about it.
Still, there's marginal returns on how much w can be increased while keeping the KDF useful, right? There's only so much magic they can do for a weak password.
The article on scrypt being a poor candidate for password storage stuck out in my memory. I haven't seen or heard of the practice of strong hash function + random salt for password storage falling out of favor, so I didn't look into using other KDFs for that purpose.
EDIT: it does seem like a trade off. This article seems to have a good explanation of it: https://crackstation.net/hashing-security.htm
If you use a key stretching hash in a web application, be aware that you
will need extra computational resources to process large volumes of
authentication requests, and that key stretching may make it easier to run
a Denial of Service (DoS) attack on your website. I still recommend using
key stretching, but with a lower iteration count. You should calculate the
iteration count based on your computational resources and the expected
maximum authentication request rate. The denial of service threat can be
eliminated by making the user solve a CAPTCHA every time they log in.
Always design your system so that the iteration count can be increased or
decreased in the future.
I do stand corrected about KDFs not being OK for password storage. Seems like they can be overkill in some situations, but in general do a good job of protecting passwords.'Strong' means different things in different contexts. SHA-256, for example, is a strong hash function: it has good preimage and collision resistance. However, a function is evaluated by its weakest point: when given a low-entropy secret (which is not the assumption in a strong hash function) and a strong hash function, the easiest attack is to bruteforce it by sampling its distribution (e.g., the most common password patterns).
Since SHA-256 is fast, bruteforce is naturally also fast: the average time-to-break for a string of entropy s is 2^(s-1) times the cost of a SHA-256 evaluation. You can build on this by, say, iterating on SHA-256:
def pwdhash(input, salt, w):
h = sha256(salt + input).digest()
for i in range(2**w):
h = sha256(h).digest()
return h
But once you get into these ad hoc schemes to make the password hash less amenable to bruteforce, you're basically reinventing PBKDF2. scrypt (and bcrypt, to some extent) is like the above, but also uses random memory accesses on a large buffer to make things harder for specialized hardware (think GPUs, FPGAs, etc).This sentence makes no sense to me. tptacek, correct me if I'm wrong, but Im pretty sure you can increase w to be as large as you want (i.e. the keyspace is large enough that youll be waiting years for a singular hash to finish before you wrap the whole keyspace). The only problem is that you force legitimate users to wait as well.
Let make up an extreme example to illustrate, you set w so that a hash takes 1 years - 1 day to compute on the clients machine. Also, they change passwords every year. (Yes, this is absurd because it means that the client can only work one day a year). Now, lets assume that he uses one of five-hundred passwords. If the attacker has 100x the compute power of the client, he will only have a 20% chance of getting the correct password. And that is with only roughly 9 bits of entropy.
(He's not arguing against using KDF-like constructions to store password authenticators).
I think that's what he meant when he said "realistically". In the reply I posted the quoted text points out that for busy auth servers, it increases potential for DoS.
FWIW, I've seen several people claim that using PBKDF2 to store passwords is something only idiots do. They never explain why, of course; they just laugh at the people who don't know better. "lol PBKDF2 idiot." I don't know where this particular thought comes from, but it's out there.
Nerds can't resist any opportunity to turn something into an editor war.