Linux RNG flaws
bugs.chromium.org
bugs.chromium.org
The main problem with the fix is that there are some userspace applications which assumed they could get cryptographic randomness super-early during system startup, and with a patched kernel, those userspace applications would block --- and in some cases, block the boot altogether. With no activity, there is no entropy to harness, and the boot scripts essentially deadlock waiting for the random pool to be initialized.
So for some hardware, and some distributions, we're getting some boot hangs that we now need to try to workaround or fix somehow. In general, the best thing to do is not to rely on cryptographic strength random number generation during early boot. Key generation should be done lazily, and deferred for as long as possible.
The other issue is that I don't get paid to work on the random driver in Linux. I've been looking for volunteers to work with the grub, syslinux, efistub, not to mention all of the various signed bootloaders used by different Android devices, but because we had a fallback mechanism there is less motivation for people to want to work on this.
This would actually be a great intern project or GSOC project, but this have been so hectic this year, personally and professionally, I didn't have the time to commit to hosting an intern or GSOC student this summer. :-(
If you clone/generate VMs from one "master" image, running virt-sysprep (or similar) is one of the steps you should do right before launching a new instance (inject a new random seed, wipe out any SSH host keys, etc.)
Even DigitalOcean missed doing this step a while back.
For paranoia you could host your own fairly easily instead of using a public one: have a machine that is publicly visible respond with random digits from its own entropy pool, have it use one or more of various methods to keep that pool topped up (local interrupt timings, a hardware RNG, ...), use request hashing with a "secret" key if you are worried about the general public draining your entropy pool or available bandwidth.
And if it can't see the entropy service for any reason (it is booting in an environment where the rest of the network is completely cut off, or perhaps the entropy service is down) then fall back to just doing what-ever is done now.
Applications can then check if the kernel is ready to do RNG and if not, handle the delay themselves.
But I guess that would break some applications so there might be another way (extra device for early boot randomness? ioctl?)
The proper fix is to change the boot scripts to not assume that the CRNG will be available directly after boot.
..for you. Other people have a different threat model where availability is ranked above RNG strength in the first minute of booting. Making security decisions not based on any sort of threat model considerations is zealotry/cargo-culting, IMO.
Above all else, the rule is that the kernel shouldn't cause userspace regressions. Something that works now must continue working later.
Atleast from what I know.
After seeding, there is no point in blocking on a lack of entropy, which is why it is recommended to use /dev/urandom in the first place.
So getrandom(2) will block until the pool is seeded, and all newly written application should use it, and existing applications should switch to it. But you still need to try to lazily generated random keys, and think very hard about whether you really need to generate cryptographic grade random numbers before the user logs in.
E.g., if on a platform with a hardware RNG then it should never be necessary to block: just read 128 bits out of the hardware RNG & generate a stream of random numbers.
On a platform without a hardware RNG, could one require a previous seed? An installer could install the seed in some persistent storage somewhere, so there's no need to do this even on boot. A VM system needs someway to atomically read the seed & then write a new seed.
On a platform without a hardware RNG and without a previous seed (i.e., the very first boot after a hand-install or something), could the system require the user to type keys until it has collected 128 bits of entropy?
I'm not certain if there's a good answer on systems with no hardware RNG, no previous seed and no input capability. Maybe CPU timing loops or somesuch?
It'd be also be nice were there a filesystem interface to getrandom(2) …
Requiring a previous seed requires a way to get access to the seed, early enough in the boot that it is available to kernel users who are trying to use randomness for address space randomization and for stack canaries. But in early boot the kernel may not be sufficiently initialized to read from persistent storage, and there are many, many bootloaders.
The reason why there is no file system interface to getrandom(2) is that a file system interface is subject to file descriptor exhaustion attacks. It was OpenBSD which designed the getentropy(2) system call, and getrandom(2) was modelled after it. Basically, getrandom(2) is getentropy(2) with an extra flags parameter added.
Typically, the crypto is needed for network communication. It can provide some randomness too...
https://lists.freedesktop.org/archives/systemd-devel/2018-Ma...
I recently updated to a 4.x series kernel and had this issue during boot. I had to interact with the VM to get it to unblock and boot.
I don’t know if that solely uses the system RNG though.
The problem here is that you would have to probe RNG stats manually as it returned readiness too early.
It is possible, however, to force it sshd-keygen, sshd, et al., to use /dev/random instead (by setting "SSH_USE_STRONG_RNG" in /etc/sysconfig/sshd to a value >= 14). In this case, would keys still potentially be at risk?
---
For a while now, I've been generating my SSH host keys at the end of the kickstart installation process (for physical hosts) with SSH_USE_STRONG_RNG=32 (in a "%post" script).
Virtual guests (KVM) still generate theirs on first boot, but they have the benefit of access to the host's /dev/random and also get a random seed "injected" into their disk image just before first boot -- although I'm not sure if that helps with this issue.
With the new patches, are the daemons benefiting from the seed ?
Additionally, CHACHA20_KEY_SIZE is a power of 2 and optimizing compilers like GCC will convert the modulus into a binary-AND.
JMG's FreeBSD 2017 RNG talk Tl;dr:
1. Boot-time entropy might leak on either UFS or ZFS; might be re-used in a byzantine ZFS environment. (Not addressed, as far as I know.)
2. Input entropy data (pre-whitening) had only about 0.18 bits of entropy per byte of input. NIST likes to see 4-6 bits per byte. (Not really an issue due to sufficient input volume.) Partially due to:
3. All the fast, high-quality sources of random entropy were accidentally disabled(!). (All the PURE_* ones, including x86 RDRAND.) Fixed in r324394.
4. Other low entropy structures were getting mixed in. Probably harmless, but reduces entropy-per-byte measure that NIST cares about. Fixed in r324372.
The earlier 2015 FreeBSD RNG issues were probably as severe as this Linux issue[2]:
> URGENT: RNG broken for last 4 months
> If you are running a current kernel r273872 or later, please upgrade your kernel to r278907 or later immediately and regenerate keys.
> I discovered an issue where the new framework code was not calling randomdev_init_reader, which means that read_random(9) was not returning good random data. read_random(9) is used by arc4random(9) which is the primary method that arc4random(3) is seeded from.
> This means most/all keys generated may be predictable and must be regenerated. This includes, but not limited to, ssh keys and keys generated by openssl. This is purely a kernel issue, and a simple kernel upgrade w/ the patch is sufficient to fix the issue.
[0]: https://www.funkthat.com/~jmg/vbsdcon_2017_ddfreebsdrng_slid...
[1]: https://www.youtube.com/watch?v=A41cDCE6pTc
[2]: https://lists.freebsd.org/pipermail/freebsd-current/2015-Feb...
Nice.
Sequencing was modified to better align expectations with actual system operation.
> == Discarded early randomness, including device randomness ==
> == RNG is treated as cryptographically safe too early ==
> == Interaction between kernel and entropy-persisting userspace is broken ==
> == No entropy is fed into NUMA CRNGs between rand_initialize() initcall and crng_init==2 ==
> == initcall can propagate entropy into primary and NUMA CRNGs while crng_init==1 ==
and > == Impact ==
> I have spent a few days attempting to figure out how bad these issues are.
> I believe that on an Intel Grass Canyon system, with RDRAND disabled,
> ASLR disabled, fast boot enabled, no connected devices, with boot on power,
> some frequency scaling options disabled, and the fan set to maximum,
> it should be possible to express the entropy in the used RDTSC samples in around
> 105 bits or less. (I'm not sure which parts of this configuration actually
> influence the amount of entropy; but ASLR certainly does influence it, since the
> one interrupt sample that is fed into the RNG before the RNG initialization
> contains an instruction pointer.)> The worst part of this (one device entropy sample being enough to move to crng_init==1) was AFAICS introduced in commit ee7998c50c26 ("random: do not ignore early device randomness"), first in v4.14.
confirmed by the fix for that bit: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
I'd only start to worry when it gets to say 70 bits or less. That's the kind of size a state actor could precompute ssh host keys for each of the initial random number states in a common distro for example.
If a state actor could compute ssh keys in 50 milliseconds on custom hardware, and they had 100,000 machines generating those keys, they could generate enough keys to have a 1 in 1 million chance for 70 bits of entropy after 20 years.
> When read during early boot time, /dev/urandom may return data prior to the entropy pool being initialized.
I see no bugs here. Just repeating what is written in the man page.
> If this is of concern in your application, use getrandom(2) or /dev/random instead.
These bugs affect getrandom too.
> Multiple callers, including sys_getrandom(..., flags=0), attempt to wait for the
> RNG to become cryptographically safe before reading from it by checking for
> crng_ready() and waiting if necessary. However, crng_ready() only checks for
> `crng_init > 0`, and `crng_init==1` does not imply that the RNG is
> cryptographically safe.
None of the reported bugs are about /dev/urandom returning data too early.
https://hn.algolia.com/?query=urandom%20manpage&sort=byPopul...
Also I think it's common knowledge that there's little entropy available during startup but developers have no control over it regardless.
I didn't know it was that bad though. That's really bad.