Red October: Go server for "two man rule" file encryption by Cloudflare
github.com
github.com
https://groups.google.com/d/msg/golang-nuts/0za-R3wVaeQ/M8_B...
Why do crypto people only want there to be a single C implementation of everything?
My advice would be don't roll your own primitives, but other than that, have at it.
Crypto is just so mind boggling hard from what I gather, there are probably hunderds of ways attacker can mess up the system.
But overall I agree - there shouldn't be a single implementation of anything ever. Crypto people need to come up with some kind of test or prover, or something that can tell programs if they are safe enough, and where they aren't. A kind of test harness or program prover.
I suppose they need to solve the halting problem, too?
I guess you should aim for good enough and not perfect.
One of the problems that is really hard to get around is the "timing attack"; by passing various values into a crypto system and seeing how long it takes to either return or error out, you can learn things, often entirely decoding a message in surprisingly short amounts of time.
To defend against this, a crypto system must make pervasive use of operations that take the same amount of time whether they succeed or fail. For instance, if you are comparing one string against another, you must compare the same amount of string regardless of the input; you can't bail out at the first difference, which is what you would normally do.
Unfortunately, as hardware gets more and more complicated this gets harder and harder to guarantee. Surprisingly small differences have been demonstrated to crack messages. And the worst part is, this is all advantage attacker; the attacker does not need to know why he's seeing timing differences, he can just figure out what they mean and exploit them. It's the defender that has to figure out (for example) that some predictive algorithm on the CPU sometimes preloads the next chunk of RAM and sometimes doesn't according to the alignment that your buffer happened to receive, depending on the (attacker-controlled) size of the first packet, and then figure out what to do about it.
The upshot of this is that crypto implemented in darned near anything other than assembler is probably flawed. C is the standard because that just isn't practical, but almost anything higher than C is dangerous. Many of the things that a high-level language consider a feature make programming proper industrial-strength crypto either difficult, or simply impossible. (Good luck preventing timing attacks in Haskell. The language and runtime fights you with every fiber of its being. And the ways in which it is fighting you are good thing the vast majority of the time... just not this one.)
It's relatively easy to produce a crypto library that emits the correct streams and decodes the incoming streams into the proper values, but that is merely the beginning of creating a truly robust crypto library; it's not even the halfway point. It's the easy part.
Go's crypto is built by people who know this stuff, so for instance: http://golang.org/pkg/crypto/subtle/ However, they do not claim it has been vetted by anybody in particular. Proving their implementation correct is very difficult, and I'd still worry about whether GC is going to bite somebody somehow in the implementation, nor is there any particular proof that they didn't miss a place they should use one of those functions, and I wouldn't bet my life the functions are 100% correct in all cases, either. (They're simple, they look good, but one stray compiler optimization and.... who knows?)
Because it's hard, and if you make a mistake, it means there is no security.
It doesn't need to be a C implementation, but crypto is very very easy to get wrong.
So never try a new implementation if you don't fully understand existing ones.
I’d rather have many imperfect-and-actively-improving implementations, than a single implementation that is widely relied upon and fails.
An example is web technology: plenty people have made plenty html/css/js mistakes in the past but there is also an overwhelming amount of evolving best practices on this topic, no doubt because of the mistakes made before. If all we do is decry crypto is hard, don't roll your own, then no one would be learning anything, and cryptography will be left to the a few select mathematicians and organisations that are referred to with acronyms.
It should be: crypto is hard, learn it and experiment with it, share your thoughts, and don't roll your own for critical applications.
Edit: clarity of argument in my example and grammar
Is it Rob Pike doing that thing that "different because I said so" thing that he and djb have been pioneering for years?
(edited to clarify that I actually don't know historically why they chose that)
It's very common on OS X. Example: http://ss64.com/osx/lipo.html
What's really striking to me is they suggest running this TLS server using command-line options instead of a configuration file. Every real daemonized network service uses a configuration file, mainly because it's much more simple & flexible to manage than whatever init script you're executing your daemon from. And anyone who thinks they should manage configuration inside a released executable (aside from defaults) is an idiot.
The major benefit is that different parts of the program can declare their flags separately, and they all get put together. When you're doing distributed flag definition like this, single-letter flags don't make much sense: you'd get collisions very quickly. Instead, a flag only has a long name, not long and short. At that point, there is no need to distinguish between - and --, so the parsers (both C++ and Go) do not.
You wrote that the parser insists on -key=value, but that's not true. Because all flags are "long names", all of these are equivalent:
-key value -key=value --key value --key=value
If you prefer one over the others, use that one.
Rob and I came from the Plan 9 (really, early Unix) world of only single-letter flags, no long names at all. That may work for small programs, but it doesn't scale very well. The GNU "short and long" is a compromise to keep backwards compatibility with the short flag world, but it ends up creating two names for many things. There's no need for that compatibility if you're not trying to recreate Unix (like GNU is). Keep it simple: one name, don't bother caring whether people type - or --.
As others have pointed out, if you need GNU getopt compatibility or even Plan 9 compatibility, it is possible to implement as a separate library. I haven't looked recently, but I've been waiting for someone to implement such a library but using the actual flag registration from the standard flag package, so that moving from one flag syntax to the other is just a matter of calling a different parser in package main. All the actual flag creation need not change.
> Despite this flexibility, we recommend using only a single form: --variable=value for non-boolean flags, and --variable/--novariable for boolean flags. This consistency will make your code more readable, and is also the format required for certain special-use cases like flagfiles.
The one remaining inconsistency between this advice and GNU's is that the latter would spell it --no-variable.
Why would you ever need to "scale" a command to more than 30 options? At that point you have a command line string that's 200 characters long, which is completely impractical for a command line. And in terms of distributed flags, who says one command wouldn't end up with the same flag name as another command?
Look at the gcc man page. It's ridiculous! They use a mix of longnames and short, and it's still confusing as hell. Why do they need this many options? Because for any given operation they want to be able to specify the overall options, language options, warning options, debugging options, optimization options, preprocessor options, assembler options, linker options, directory options, machine-dependent options, platform options, code gen options, etc. As if all of those options would relate to every single GCC operation!
They seem to have a mandate that input from ARGV is the only way to determine how to run their program. And it works for them; using autoconf and Makefiles you can configure the options each command will use in a very fine-grained way. But the end result is that for 99% of these cases, you have to use autoconf and Makefiles to make sense of anything! So what the hell is the point of stuffing them all in ARGV if a human isn't even going to use them anyway? They could have just as easily set environment variables and been done with it (lord knows why; probably some 30-year-old backwards compatibility issue)
I prefer the 'cvs' model of command-line processing. The creators knew they wanted one command which would let them do everything, but they also knew that there would be separate operations with their own options. So you specify a command, and then you specify your options. In this way, you know -a for 'ls' is probably going to be different than -a for 'diff' - but you can also look up these options incredibly easy using 'cvs --help command'. Now instead of having a soup of options to browse through, you know exactly what options relate to the specific operation you want to do.
The program's execution and options become easier to understand using this option hierarchy. It also distributes well because your commands have unique names. And no more looking up the 1000-page manual to find the one option you wanted.
> -key value -key=value --key value --key=value
> If you prefer one over the others, use that one.
Thank you, I had no idea! I support this design choice now.
http://blog.cloudflare.com/red-october-cloudflares-open-sour...
Because it uses existing applications there's virtually no risk of the crypto implementation having bugs (or only openssl's bugs, anyway). And because it's so simple you can add whatever functionality you want to fit your system. For example, you can modify it to require a passphrase on the keys; I left it out for simplicity.
https://github.com/psypete/public-bin/blob/public-bin/src/se...
This image for example: http://blog.cloudflare.com/static/images/cryptography2.png stores identical plaintexts under different encryptions.
> Red October is backed by a database of accounts stored on
> disk in a portable password vault. The server never stores
> the account password there, only a salted hash of the
> password for each account. For each user, the server
> creates an RSA key pair and encrypts the private key with
> a key derived from the password and a randomly generated
> salt using a secure derivation function.
Breaking this down:1. Create database 2. Create user 3. Store salted/hashed user password 4. Create pub/priv keypair for user 5. Encrypt keypair using password and a new salt
If you get the vault, you need the user's password to decrypt their keypair. Trying to crack their password is really hard (they use scrypt). Trying to crack an encrypted keypair is really hard because they also use scrypt (i assume?). Even if it is cracked, all the other encrypted keypairs have a different salt, so even one cracked still leaves all the others to be cracked again.
So even if the same plaintext data is encrypted by multiple people, it'll be substantially different each time because of all the different keypairs and salts.
The attacker knows that identical processes are used to encrypt that data-key. Even with multiple keys and nested encryptions, it seems that at the very least that the reduced entropy of search space around the data-key can be used to optimize a brute-force search. This would be further compounded if at least one user's key-encryption-key were known.
I'm probably missing something important, but if I've learned anything from the cryptopals challenges its that you don't want to reduce your entropy in any way.
Would a simple solution be to add a random salt at each level of the nested key encryptions?
"No data is leaked provided the encryption scheme is IND-CPA. AES CBC with a random IV is believed to meet that security definition (within a bound of ~2^64 separate encryptions with the same key and assuming the AES PRP is itself secure)."
http://en.wikipedia.org/wiki/Ciphertext_indistinguishability
Just prior to that, they use their keys to jointly open a safe containing the file with their orders (analog to decryption?). Maybe that's what the name is alluding to.