Myths about /dev/urandom
2uo.de
2uo.de
They know their crypto requires good entropy. GPG should timeout or give a warning/option to shoot yourself in the foot in case you wanted to. But that's UX, not crypto.
As the article says, it's a screw up to not have each VM instance seeded separately. You need that to not muck the whole thing up. In the case of generating keys for permanent use, you want to fail in that case, at least until some entropy shows up, as it did.
I would agree with the article that the right way to fix the problem is to seed each instance of the image on startup, but that will also avoid you having a problem with it blocking.
That's a special case though, where you aren't in the middle of a cryptographic handshake, you don't have real time constraints, and the fix to the real problem will also mean there is no problem. Don't use it for a network service's source of randomness.
You do not need /dev/random for key generation.
I think blocking and not returning until you have entropy is a reasonable failure behaviour for gpg in the key generation process.
It'd be nice to report something and maybe hint to the user that you are waiting for just a modicum of entropy to show up, but at least it isn't presenting a key that is entirely predictable (and worse still, the SAME as any other instance of that VM image!!!) to anyone with access to the original VM.
The bug the article is referring to is that a lot of security systems will block reading /dev/random when in fact /dev/urandom will provide a securely unpredictable sequence of data with no statistically likelihood of another system producing the same sequence. It's particularly bad for the case where timeliness is an important part of the protocol (which is largely a given for anything around say... a TCP connection). That's a silly design flaw.
But, if that's the case, then the entire thesis behind, "Use /dev/urandom" is incorrect. We can't rely on /dev/urandom, because it might not generate sufficiently random data. /dev/random may block, but at least it won't provide insecure sequences of data.
This is kind of annoying, because I was hoping that just using /dev/urandom was sufficient, but apparently there are times when /dev/random is the correct choice, right?
That was the point of the blog post: if you are using /dev/random as input in to a cryptographic protocol that will stretch that randomness over to many more bytes, WTF do you think /dev/urandom is already doing?
What /dev/urandom might fail to do, and this primarily applies to your specific case of a VM just as it is first launching and setting things up, is generate unpredictable data, and most terribly, it might generate duplicate data, for certain specific use cases where that would just be awful.
I would agree that you got the gist right though: /dev/urandom is usually the right choice, but when it is not, /dev/random is indeed needed. Most people misunderstand the dichotomy as "random" and "guaranteed random", which leads to very, very unfortunate outcomes. Other people misunderstand what they are doing cryptographically and somehow think that some cryptographic algorithm that uses a small key to encrypt vast amounts of data shouldn't have any concerns about insecure, non-random cyclic behaviour, but oddly take a jaundiced eye to /dev/urandom. It basically amounts to "I think you can securely generate a random data stream from a small source of entropy... as long as that belief isn't embodied in something called urandom".
Again, if you don't know the right choice, you should pass the baton to someone who does, because even if you make the right choice, odds favour you doing something terrible that compromises security.
That was not something that I was aware of, thanks.
"Daniel Kahn Gillmor observed that GnuPG reads 300 bytes from /dev/random when it generates a long-term key, which, he observed, is a lot given /dev/random's limited entropy . Werner explained that GnuPG has always done this. In particular, GnuPG maintains a 600-byte persistent seed file and every time a key is generated it stirs in an additional 300 bytes. Daniel pointed out an interesting blog post by DJB explaining that a proper CSPRNG should never need more than about 32 bytes of entropy. Peter Gutmann chimed in and noted that a 2048-bit RSA key needs about about 103 bits of entropy and a 4096-bit RSA key needs about 142 bits, but, in practice, 128-bits is enough. Based on this, Werner proposed a patch for Libgcrypt that reduces the amount of seeding to just 128-bits."[^1]
On a related not why are you generating keys on a remote vm? Its probably not fair to say that gpg "failed to work (literally would not do anything)." It was doing something, gnupg was waiting for more entropy. Needing immediate access to cryptographic keys that you just generated on a recently spun up remote VM is kind of a strange use case?
Re: "Why are you generating keys on a remote VM" - prior to this, it hadn't occured to me I couldn't generate a gpg key on linode/digital ocean, VMs. I realize now that keys should be generated on local laptops (or such), and copied up.
Re: "Fair to say failed to work" - It just sat their for 90+ minutes - I spent a couple hours researching, and found a bug (highly voted on) that other people had run into the same issue. But, honestly - don't you think that gpg just hanging for 90+ minutes for something like generating a 2048 bit RSA key should be considered, "failing to work?" - I realize under the covers (now) what was happening - but 99% of the naive gpg using population would just give up in the same scenario instead of trying to debug it.
Hard to get that kind of thing right, but fundamentally it did stop you from making exactly the kind of terrible mistake that I was talking about. ;-)
http://man7.org/linux/man-pages/man2/getrandom.2.html
With the flags set to zero, it works like the getentropy(2) system call in OpenBSD. In fact, code that uses getentropy(buf, buflen) can be trivially ported to Linux as getrandom(buf, buflen, 0).
1) Why does the Linux system call have flags at all?
2) Why does the Linux system call have a different name?
I.e. why not clone the OpenBSD system call exactly with the same name and semantics and not give the userland any flags to shoot self in the foot with?
(you certainly don't have to agree with that, currently all the flags can do is draw from /dev/random instead of /dev/urandom, and do so without blocking, but that's the reasoning behind the interface change and thus the name change)
getrandom() showed up in 3.17, and the latest CentOS distribution ships with 3.10. And then you've got older deployments.
You could play with autoconf hell of detecting getrandom() and compiling for it if needed, but I couldn't imagine it to be worth the effort at the moment if you're going to be maintaining a branch of fallback code anyway.
OpenBSD and NetBSD feature a ChaCha-based arc4random; FreeBSD and libbsd still seem stuck with RC4-based arc4random[1]. An equivalent for that is sorely missing on Linux. /dev/urandom requires messing with file descriptors, which you may run out of and may require error handling, plus all kinds of security precautions to make sure you're actually looking at the right /dev/urandom[2].
[1] https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=182610 and https://bugs.freedesktop.org/show_bug.cgi?id=85827
[2] http://insanecoding.blogspot.ch/2014/05/a-good-idea-with-bad...
It's unfortunate that it's so recent, though.
If there's anything I've learned by studying cryptography, it's that the average person needs to invest significant resources to become a pseudo-expert in cryptography - all before they can make any meaningful decision.
Applied Cryptography is not a terribly good starting point for cryptographic understanding these days. That's just another facet of difficulty: where does one even start? Do you want a theoretical baseline understanding? Do you want the high-level, quick-and-dirty overview which only gives a summary? (That'll make it hard to make real, informed decisions...)
If we really want to make cryptography accessible to many developers, the best solution is for the cryptographic community to make our libraries and interfaces better. At the same time, there comes a point where a developer has to stop and say "this is beyond what is standardized in ${widely used library}; we need to hire a cryptographer." In that sense, we need to instill a better anti-crypto ethic.
About starting point, this is incredibly sad but true: there is no Applied Cryptography of 2015. The book is at this point in some way outdated and no replacement exists, however what you can do is to read it, and then to read the documents that there are around to get updated information. Also there are now the online courses on cryptography that really help. This may look like an overkill, but at this point crypto is everywhere and is the foundation of most things secure, so it is a requirement of everybody involved with computer security.
All of these influence how the secret should be bundled up and sent, and it takes more than a library to pick the appropriate method.
There's no way around the fact that security is complex. If it's too time consuming for the average developer, then they shouldn't be doing it at all.
I think what you're getting at is a specific technology won't save you, it's the security mindset that's needed. Convincing programmers there really are bad people who want to pick apart your systems is the problem. Once their convinced, once they take security seriously, they'll do better.
They might start out doing a terrible job, but with the security mindset, they'll improve. They'll seek out problems and solve them. Rather than pretending it's not an issue, or blindly apply security secret sauce like "use bcrypt"
Do you perhaps mean Cliff Stoll?
[1]: https://github.com/andrewrk/genesis/blob/0d545d692110d33068d...
The excerpt below is from https://www.nsa.gov/ia/_files/factsheets/I43V_Slick_Sheets/S... (which in turn also references https://www.nsa.gov/ia/_files/factsheets/I43V_Slick_Sheets/S... )
Unix-like Platforms (e.g. Linux, Android, and Mac OS X):
Application developers should use the fread function to read random bytes from /dev/random for cryptographic RNG services. Because /dev/random is a blocking device, /dev/random may cause unacceptable delays, in which case application developers may prefer to implement a DRBG using /dev/random as a conditioned seed.
Application developers should use the “Random Number Generators: Introduction for Operating System Developers” guidance in developing this solution. If /dev/random still produces unacceptable delays, developers should use /dev/urandom which is a non-blocking device, but only with a number of additional assurances:
- The entropy pool used by /dev/urandom must be saved between reboots. - The Linux operating system must have estimated that the entropy pool contained the appropriate security strength entropy at some point before calling /dev/urandom. The current pool estimate can be read from /proc/sys/kernel/random/entropy_avail.
At most 2^80 bytes may be read from /dev/urandom before the developer must ensure that new entropy was added to the pool.
And I wouldn't put it past them to have an attack (at least theorectical) that exploits this
So if they write that there's a vulnerability after reading 2^80 bytes - that's great! We're secure. If they write that you must ensure to do something after 2^80 bytes - that's complete bullshit.
However, remember when attacks to 3DES, MD5 were only theoretical?
Also, you may not even need to read 2^80 bytes, there might be a (future) vulnerability that allows you to shortcut this.
Now, if someone found a lower limit based on exploiting some weakness in the random number generation, the analogy with 3DES and MD5 would make more sense.
If you can break reality, all bets are off though and trying to defend against attacks that break reality in the future are impossible.
But there seems to be smarter attacks (ref 21) https://en.wikipedia.org/wiki/Triple_DES#Security
My point is that even if today some sizes and lengths seem only of theoretical concern, tomorrow there might be a vulnerability, a new approach to the problem, or even natural technological evolution that might turn it into a practical attack (even when it seems impossible today)
[1] http://www.theguardian.com/business/2009/may/18/digital-cont...
(Except for saving the entropy pool between reboots, which can be useful, afaik.)
Wait, what? dev/urandom is a DRBG seeded the same way dev/random is. This doesn't seem to make sense.
You're very unlikely to be continually reseeding your own DBRG with new entropy, so it will be less secure than /dev/urandom, which is.
Suspiciously bad advice, there, from what I can see.
But yes, that's not common. (more info: http://wiki.openssl.org/index.php/Random_Numbers)
No. Doing random in userland is just wrong. If your program has access to /dev/urandom, use it. If not, use arc4random().
> “Random Number Generators: Introduction for Operating System Developers”
Or look how OpenBSD does it. (getentropy(), arc4random(), the subsystem)
> The entropy pool used by /dev/urandom must be saved between reboots.
OpenBSD does this, and more. The bootloader basically seeds the kernel with old entropy from before the reboot.
Every single Linux distro does this.
(Ego is the reason for a bunch of other security misfeatures in Linux. Securelevel comes to mind, where Linus explicitly said after fifteen years, okay, this was in fact the right model all along, you can merge it, just call it something other than securelevel so I don't have to eat my words about securelevel being a mistake.)
rm /dev/random && mknod -m 444 /dev/random c 1 9BSD's implementation seems sensible -- block the first time and then don't block again. Or potentially block if you haven't reseeded in X amount of time (X being configurable), to defeat the situation where someone figures out your seed. Although, I have a tough time understanding how someone will figure out your seed without completely compromising your system anyway... but systems are supposed to be robust against attacks that one can't immediately imagine ;-)
I know that a PRNG is predictable if you know all the input variables, and the code for it is publicly available, but has anybody in practice been able to exploit that?
EDIT: that's an honest question. I'd like to read a paper about that.
[1]: http://www.linuxplanet.com/news/linux-4.2-released-improving...
Fun thing is, if you pass "/dev/urandom" to this parameter, Java will read /dev/random anyway. May be that was a wise decision 20 years ago.
For a concrete example, you start out with zero bits in the pool. /dev/urandom will produce a predictable stream of bits. This is extremely bad. /dev/random will block. this is good.
Now you add 1024 bits to the pool. Both /dev/random and /dev/urandom will produce good numbers.
Now let's say you read 1024 bits from /dev/random. This will reduce the entropy counter back to zero. If you then try to read another 1024 bits from /dev/random, it will block.
But this is nonsense! Those 1024 bits you added before aren't depleted just because you pulled 1024 from /dev/random! It is perfectly safe to proceed generating more numbers at this point (or at least it's as safe as it was before), but /dev/random refuses to.
Many systems don't add entropy to the pool very quickly so it's entirely possible to "deplete" it in this way. Then your code using /dev/random wedges and you have a problem. However, few systems are ever in a state where they have zero entropy. Thus the advice to use /dev/urandom.
Ideally, you'd want a device which blocks if and only if the entropy pool hasn't been properly seeded yet, but which never blocks again after that point. Apparently this is what the BSDs do, but Linux doesn't have one.
An observer of the produced random numbers can potentially deduce the next numbers from the first 1024 random numbers. This is the reason why /dev/random requires more randomness added -- to prevent the guessing of the next number.
By definition, a cryptographically secure pseudorandom number generator cannot be predicted like that by a computationally bounded attacker.
Thus if any attacker could deduce the next number from /dev/random by observing the numbers before, the algorithms they adopt is fundamentally wrong, and nothing could the save the security in that case.
(That's why I've chosen this structure for the essay, and have included anchors)
Think of a CSPRNG almost exactly the way you would a stream cipher --- that's more or less all a CSPRNG is. Imagine you'd intercepted the ciphertext of a stream cipher and that you knew the first 1024 plaintext bytes, because of a file header or because it contained a message you sent, or something like that. Could you XOR out the known plaintext, recover the first 1024 bytes of keystream, and use it to predict the next 1024 bytes of keystream? If so, you'd have demolished the whole stream cipher. Proceed immediately to your nearest crypto conference; you'll be famous.
Modern CSPRNGs, and Linux's, work on the same principle. They use the same mechanisms as a stream cipher (you can even turn a DRBG into a stream cipher). The only real difference is that you select keys for a stream cipher, and you use a feed of entropy as the key/rekey for a CSPRNG.
It's facts like this that make the Linux man page so maddening, with its weird reference to attacks "not in the unclassified literature".
The getrandom() call added in Linux 3.17 does this. By default it reads from urandom, but it blocks until it estimates that urandom has been seeded with "a sufficient number of bits", which seems to be around 128 bits of randomness.
http://man7.org/linux/man-pages/man2/getrandom.2.html
When available getrandom() should probably be preferred over both /dev/random and /dev/urandom.
A further complication - you don't actually add 1024 bits to the pool. Instead, you add N bits to the pool (N > 1024), estimate its entropy as 1024, and increase the overall entropy estimate by that much.
Estimates are always pessimistic, because it's better to be cautious and under-estimate the entropy than the alternative. So you might have 2048 bits or more of real entropy, but that can still be "depleted" by a request for 1024 bits.
The paper that describes Fortuna [0] goes into greater detail on this. Yarrow, the CSPRNG that preceded Fortuna, attempted entropy estimation. Fortuna rejects entropy estimation, and tries to build a CSPRNG that is secure without needing estimates:
"...making any kind of estimate of the amount of entropy is extremely difficult, if not impossible. It depends heavily on how much the attacker knows or can know, but that information is not available to the developers during the design phase. This is Yarrow’s main problem. It tries to measure the entropy of a source using an entropy estimator, and such an estimator is impossible to get right for all situations."
"Fortuna solves the problem of having to define entropy estimators by getting rid of them."
What really annoys me is that a Fortuna-based CSPRNG was contributed to Linux eleven years ago this month, but died on the vine, IIRC because the Linux kernel team were so fond of the entropy-estimation approach.