Quality entropy isn't a hard ask, at least not for anything typically running OpenSSH or other server software. Intel has RDRAND, and even where AMD's RDRAND is broken their PSP coprocessors provide an entropy function. Similarly, NICs and other controllers also often come with RNGs. RNGs abound on modern embedded systems, actually, it's just that nobody has the full-time job of plugging them into the kernel's PRNG pool. It's not a full-time job for any OpenBSD developer, either, but they do seem to do a better job of this than on Linux; often times the only feature supported of a miscellaneous system component is the RNG.
It's the APIs that are broken. getrandom shouldn't block, period. A system only needs 16-32 bytes of pure hardware randomness for strong security. That's it! You either have it shortly upon boot, or you don't. If you don't, you're screwed anyhow, so why block? If you have 32 bytes of good entropy, all the entropy accounting mumbo-jumbo is pointless.
I do take issue with the author's complaint that systemd shouldn't have its own user-space PRNG. BSD systems have arc4random() as part of their libc, which is seeded from the kernel pool. This is very convenient; developers shouldn't have to think twice about calling into a PRNG for a 32-bit number, but if acquiring that requires a syscall they do think twice and often screw things up. Not to mention that Linux getrandom's default block semantics is broken by design.[1] Until glibc, musl libc, and other Linux runtimes wisen up and add arc4random, it's hard to blame projects like systemd for including their own PRNG.
[1] I realize that getrandom has an option to not block, but it's only function is to cast suspicion on itself when in fact the only thing that deserves suspicion are entropy guesstimators.
128 bits is enough to feed into a stream cipher that will generate a lot of random bits:
> The original version of this random number generator used the RC4 (also known as ARC4) algorithm. In OpenBSD 5.5 it was replaced with the ChaCha20 cipher, and it may be replaced again in the future as cryptographic techniques advance.
* https://man.openbsd.org/arc4random.3
* https://security.stackexchange.com/questions/85601/
Re-feed/-stir every so often so that past entropy state can't be used to compromise things in the future.
Correct.
> You either have it shortly upon boot, or you don't.
Not correct. On virtualized systems, with minimal interrupts, there is frequently not enough entropy available. This comment seems to reflect a poor understanding of the getrandom(2) system call and its history; part of the reason for implementing getrandom was in order to provide this new behavior of no longer blocking once enough entropy has been collected for a secure CSPRNG initialization. Linux needs this "entropy accounting mumbo jumbo" because it doesn't have a standard mechanism for persisting random seeds across boots; OpenBSD doesn't need it because the bootloader and kernel are tightly integrated, so they can easily implement this feature.
Entropy accounting can't fix the underling problems here, all they do is obscure and confuse. Hypothetically and in a very technical sense, they can be useful and even necessary. But in practice they simply have no place outside of the actual hardware-based entropy generating devices. If you can't quickly seed yourself with 32[1] bytes of cryptographically strong entropy, then you're screwed, period. Blocking doesn't improve security, it just induces people and developers to implement awkward workarounds with the net effect of drastically reducing security.
When Linux added the getrandom syscall they should have dropped support for blocking. It was patterned after OpenBSD getentropy, which doesn't block; neither did Linux' long deprecated sysctl() random UUID mechanism that many programs once relied upon (like Tor). But they, and Ts'o in particular, seem unable to resist the siren call of entropy guesstimation.
[1] Even 16 bytes is enough to seed the system pool for an indefinite period, at least relative to a system without a strong hardware RNG. And that's the point. There are reasons for why a component might need ongoing sources of strong entropy, but if a system can't even provide 16-32 bytes at boot then those arguments are purely hypothetical because there clearly aren't sources of strong entropy available, anyhow. But if those sources are available, then it's ridiculous to think they can be "depleted" as a practical matter. If the CSPRNG pooling functions are broken, then all modern cryptography is broken, so you gain nothing with the convoluted semantics of Linux's traditional /dev/random machinations.
I'm not a fan of trusting closed hardware for randomness.
Fortunately in practice the situation isn't so dire. There are usually multiple hardware RNGs on a system. They may be all untrusted, but I'll trust them together before I'll ever trust the strength (and persisting correctness) of guesstimators. Ironically, entropy guesstimators are predicated on the very notion that the hardware is benign. Malicious hardware could implement subtle timing patterns in interrupts, much like what they might do for an actual RNG (e.g. use an AES encryption function to "randomize" their visible behavior). And in any event, you can still mix various system timing events into your pool without pretending you can quantitatively and reliably know their entropic contribution.
For example, the server could have an RSA private key, and the clients have the corresponding public key. A client could generate a full RSA block of random data, and encrypt that with the public key, and send the result to the server. The server could recover the random data and use that to initialize a CSPRNG.
This CSPRNG could be user mode code running in the ssh system, and only used for randomness needed for that particular connection, so that if a client supplies poor random data it only weakens that client's connection.
don't blame the linux kernel, we're just waiting on the dbus interface, gnome3 applet, and systemd binary + unit files before we integrate /etc/grub.conf.d/entropy.d/randomseed.d/grub-entropy-randomseed.conf and it's good buddy systemd-enable-grub-conf-d-entropy-d-randomseed-d.target into the Loobuntitis GNU/Linux Kubernetes CloudPaaSOS 2019.07.18 'doofy nerdguy' LTS Server release.
OpenBSD guys are clearly antiquated and using old modes of thinking with such a simplistic and neanderthal like design of having a system wide library installed in a single location without a dynamic json configurable runtime rest api endpoint.
https://fedoraproject.org/wiki/Features/Virtio_RNG
For embedded systems, many have inbuilt RNGs, have a look at your SoC docs.
edit: but i guess that's on purpose to handle disk imaging
Per some of the comments in the Debian bug(s): if you tell users and devs simply "you deal with it", you will get a multitude of ad hoc, half-baked solutions that may nor may not be secure. You will have dozens of people trying to re-invent the wheel.
The whole point of things like /dev/urandom and getentropy() is that the problem is solved once properly, and then you allow everyone to benefit. By having the above solutions not-work, we are basically going back to the days of where they did not exist--in which case what was the point of creating them?
The BSDs seem to not have a problem with this, so I have no idea why the Linux folks can't seem to get their act together.
> Key generation is already slow anyway.
Define "slow". Key generation was slow in the early 1990s when I first started using PGP (nee GPG) back in the day. It is generally not-slow nowadays IMHO--as in, it takes less than ten minutes to find p and q for RSA 1024 (nevermind 2048+).
To a first approximation:
* SHA256( cat /var/lib/randomseed || ifconfig || date -u ) > /dev/urandom
will get any decent stream-cipher-based PRNG going.
Each machine will have a unique MAC address, so that's 48 bits right off the bat, even if randomseed is identical and every machine is booted at the exact same second.
If the system is not communicating with other systems... what attack tree are you actually worried about if it is inaccessible?
I would also be curious to know which virtualization environments generate duplicate MAC addresses over (say) dozens of systems?
qemu uses the same MAC address by default.
Are you telling me there is not even, say, 8 bits' worth of entropy credit in there? (Especially when hashed with other data?)
Remember: the context of this discussion is VMs (and not embedded systems without a battery-backed RTC).
Persistent state is hard. For example, some systems run with read-only filesystem.
Edit: Replaced "host keys" with "cryptographic state".
Read and parse the config, fork, and lazily build an entropy pool.
Let the ssh connections or key generations hang if there isn’t enough entropy yet, but don’t get in the way of booting.
This is two bugs: one in systemd, and one in OpenSSH.