Similar story for systemd of course, nobody would expect them to learn their lesson. How nice of an init system to keep your system dead because it apparently needs CPRNG generated job ids of all things.
So it does not really matter if it was intentional or not: regardless of the root cause of this bug, it should now be considered that rdrand should never be trusted (unless you truly don't care that your systems risk to randomly break or you risk to randomly generate compromised keys) and its mere usage (except maybe indirectly in mitigated ways designed by experts and audited for years) considered a very likely security vulnerability.
When PulseAudio introduces random crackles and scratches in sound output, they claim, that I have a misbehaving sound card. It does not "misbehave", when I use the same applications with ALSA, but whatever.
When rdrand fails to produce high-quality entropy — again and again, over and over — Poettering complains about buggy CPUs.
I hope, that this particular group never ends up working on network hardware.
I mean RDRAND is a great example here, if you never use it, you never encounter these bugs. It's not impossible that PulseAudio tries to play higher quality audio or simply uses the device differently.
> I hope, that this particular group never ends up working on network hardware.
I'd kinda love it, probably quite a few fun hardware bugs will be found.
https://devblogs.microsoft.com/oldnewthing/20050512-48/?p=35...
Pulseaudio mixes sound from several applications, unlike Alsa. Alsa just configures your soundcard for that one stream, Pulseaudio has no such luxury.
In order to mix relatively small buffers (because latency), it needs to do timing. If your soundcard has problem with timing (and many do!), you will get crackling. The good news is, that the problem with timing can be only at specific frequencies; in the past, I had a computer with Nvidia Ion and it did this with 44,1kHz audio. With 48kHz, everything was OK. So fixing the output at 48kHz by forcing Pulse to mix at this frequency fixed the crackling.
Dmix for alsa was buggy, because just like PA, it didn't know about quirks of the hardware out in the field. And it stayed buggy, because nobody used it in return. Yes, PA was forced on users, but in the result the quirks were found, implemented and today, we all are better off. I didn't have any problem with PA since 2012 or so, but quite enjoyed the features it brought to the desktop.
https://github.com/systemd/systemd/blob/bcac754d66374782a85a...
int rdrand(unsigned long *ret) {
/* So, you are a "security researcher", and you wonder why we bother with using raw RDRAND here,
* instead of sticking to /dev/urandom or getrandom()?
*
* Here's why: early boot. On Linux, during early boot the random pool that backs /dev/urandom and
* getrandom() is generally not initialized yet. It is very common that initialization of the random
* pool takes a longer time (up to many minutes), in particular on embedded devices that have no
* explicit hardware random generator, as well as in virtualized environments such as major cloud
* installations that do not provide virtio-rng or a similar mechanism.
*
* In such an environment using getrandom() synchronously means we'd block the entire system boot-up
* until the pool is initialized, i.e. *very* long. Using getrandom() asynchronously (GRND_NONBLOCK)
* would mean acquiring randomness during early boot would simply fail. Using /dev/urandom would mean
* generating many kmsg log messages about our use of it before the random pool is properly
* initialized. Neither of these outcomes is desirable.
*
* Thus, for very specific purposes we use RDRAND instead of either of these three options. RDRAND
* provides us quickly and relatively reliably with random values, without having to delay boot,
* without triggering warning messages in kmsg.
*
* Note that we use RDRAND only under very specific circumstances, when the requirements on the
* quality of the returned entropy permit it. Specifically, here are some cases where we *do* use
* RDRAND:
*
* • UUID generation: UUIDs are supposed to be universally unique but are not cryptographic
* key material. The quality and trust level of RDRAND should hence be OK: UUIDs should be
* generated in a way that is reliably unique, but they do not require ultimate trust into
* the entropy generator. systemd generates a number of UUIDs during early boot, including
* 'invocation IDs' for every unit spawned that identify the specific invocation of the
* service globally, and a number of others. Other alternatives for generating these UUIDs
* have been considered, but don't really work: for example, hashing uuids from a local
* system identifier combined with a counter falls flat because during early boot disk
* storage is not yet available (think: initrd) and thus a system-specific ID cannot be
* stored or retrieved yet.
*
* • Hash table seed generation: systemd uses many hash tables internally. Hash tables are
* generally assumed to have O(1) access complexity, but can deteriorate to prohibitive
* O(n) access complexity if an attacker manages to trigger a large number of hash
* collisions. Thus, systemd (as any software employing hash tables should) uses seeded
* hash functions for its hash tables, with a seed generated randomly. The hash tables
* systemd employs watch the fill level closely and reseed if necessary. This allows use of
* a low quality RNG initially, as long as it improves should a hash table be under attack:
* the attacker after all needs to trigger many collisions to exploit it for the purpose
* of DoS, but if doing so improves the seed the attack surface is reduced as the attack
* takes place.
*
* Some cases where we do NOT use RDRAND are:
*
* • Generation of cryptographic key material
*
* • Generation of cryptographic salt values
*
* This function returns:
*
* -EOPNOTSUPP → RDRAND is not available on this system
* -EAGAIN → The operation failed this time, but is likely to work if you try again a few
* times
* -EUCLEAN → We got some random value, but it looked strange, so we refused using it.
* This failure might or might not be temporary.
*/So, a PRNG with a semi-decent (not perfect) entropy pool is required at boot time, and systemd needs to run in environments where seeding that pool in software could take on the order of minutes. This is why RDRAND is used to seed the pool during boot.
https://github.com/systemd/systemd/blob/61bd7d1ed595a98e5fbf...
2. Use a tree instead of a hash-table; or just use RDRAND and hope you aren't presented with malicious inputs.
The comment in the code[1] about RDRAND being sufficient for the use-case of UUIDs is empirically false. To continue to rely upon it is to insist that the rest of the world conform to your beliefs, rather than to adjust your beliefs to match the rest of the world.
1: https://github.com/systemd/systemd/blob/bcac754d66374782a85a...
Ipse dixit. Can you provide an explanation?
2. https://arstechnica.com/gadgets/2019/10/how-a-months-old-amd...
3. https://linuxreviews.org/RDRAND_stops_returning_random_value...
4. There are also comments elsewhere on this page that older Intel RDRAND implementations had similar issues, but they did not cite their source.
RDRAND is rather slow, especially with the SRBDS microcode update. One instruction takes 1,200 ns on my system, that's 6.6 MB/s. A software implementation of a CSPRNG will do hundreds of MB/s, the general purpose PRNGs do several GB/s.
I recall discussions to maybe rely on something like jitter entropy when the pool is empty and getentropy() is called.
This ended up being implemented in 5.4 kernel: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
Systemd is literally the only software, that has this problem. I am not aware of any other software, that uses rdrand and expects high-quality cryptography-grade randomness. Precisely, because is does not work. Intel CPUs used to have very similar issues with rdrand and so did AMD. Furthermore, CPU implementation of rdrand is a very attractive targets for state backdoors, so most sensible developers either follow the "GNUPG way" (ask for randomness from user) or simply read from /dev/random.
The rdrand instruction is great for games, because it allows to make white noise without complex algorithms and system call overhead. It is also handy for few situations, like interrupt handlers, when you needs to create some semi-random value without using stack space or calling into outside code. Unfortunately, when it was introduced, rdrand was documented to generate "cryptographically strong random numbers" (did Intel developers ever knew, what that means?) Of course, most actual cryptography experts didn't buy into that. As a consequence, and because multi-platform software needed to have it's own RNG anyway, the instruction remained largely unused for actual cryptography. Unused = untested and sometimes broken. If I were in systemd developer's place, I would not hinge bootability of my systems on something like that.
I know it's in vogue to rip on systemd, but let's at least try to be fair.
Eh. You're looking at hundreds of cycles per use. That makes it slower than a secure software RNG, let alone an insecure one.
Please seek out and understand the reasons behind calling RDRAND in systemd before making statements like these. "High-quality crytography-grade randomness" is explicitly not required for the purposes of the PRNG at boot-time, which include UUID generation and seeding hash tables.
https://github.com/systemd/systemd/blob/61bd7d1ed595a98e5fbf...
- The clock might not be initialized; the justification for using rdrand specifically calls out embedded systems; embedded systems are also likely to have clocks that reset to some date every boot and get fixes as part of the boot process that happens later.
- The "node ID" (nominally, a MAC address) udev has not started yet and therefore has not initialized the network cards yet (in fact, if you configure it to randomize your MAC address (for privacy), the random_bytes() routine being discussed is the one that it will use to generate the MAC!) The filesystem hasn't yet been mounted, it can't use /etc/machine-id or anything like that either.
* RDRAND * Seed from previous boot * Jitter * Result of every sensor (for example temperature) attached to the machine * Seed from hypervisor (if we are in a VM) * Time & Date * Any information about the physical computer we are on (attached hardware)