A months-old AMD microcode bug destroyed my weekend
arstechnica.com
arstechnica.com
That's just... not true. Is the article wrong, or was the rdrand issue actually bios-specific rather than in CPU microcode?
Edit: Whoever downvoted this - ucode update is supported by both Intel and AMD for quite a while. You certainly can just download it: https://github.com/platomav/CPUMicrocodes/tree/master/AMD
On Ubuntu, he'd want to use the amd64-microcode package: https://launchpad.net/ubuntu/+source/amd64-microcode
Which a live boot, you can install packages such as microcode updates.
You'll need to compile a kernel which does not depend on the RDRAND instruction as a source of randomness.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...
Although in the author's case, he seems to be waiting on a BIOS fix because the microcode fix doesn't actually work. In the updates he mentions being told "...that the amd64-microcode package... would fix the issue. This, too, is not the case; the amd64-microcode and intel-microcode packages are both installed by default on all Ubuntu 19.10 systems, including the one I'm experiencing the RDRAND failure on."
It's good news that microcode fixes generally don't demand BIOS updates, but in practice this one seems to be getting fixed that way before it's fixed by the manufacturer.
That's as well my understanding of the situation => does anybody have more infos?
I was really about to upgrade from an i7-8700 to a Ryzen 3700X but after reading this I froze and now I feel totally insecure... .
EDIT: was answered by "iforgotpassword"
①: https://www.gigabyte.com/Motherboard/X570-GAMING-X-rev-10/su...
Can anyone comment why the AMD microcode update doesn't seem to actually fix this bug? I've seen plenty of people commenting that you don't need a bios update and that it should, however the author seems to have tried that.
Also, the Ryzen 3 seems to be picky in dealing with RAM chips. I couldn't boot my computer with all slots in my motherboard filled. I ended up having to dial down the speed of my RAM in the BIOS as well as manually set some other settings.
Also, certain motherboards require an updated BIOS before it can boot off a Ryzen 3. In my case, I had the store I bought the chip from install the chip into the mobo so that they could demonstrate that it would boot up and I wouldn't have to return anything. It took them hours to figure this out and I had to tell them how to fix the problem after arriving because I had previously run across mention of the problem. The fix was to install a Ryzen 2, enter the BIOS, install the upgrade and then go ahead with installing the Ryzen 3 chip.
What a PITA that was. It put me off building my own system. Next time, I think I'll get a Dell or something.
I understand you can't just get any memory. The memory I purchased was reasonably within the required specs and a solid brand. The memory wasn't on the "tested list" but on a relatively new chip that's a small list and in the Philippines, I had a limited choice. This is also a wide problem as I found loads of people who had the same issues. I was able to fix it though, and I have my original RAM installed. If there's a secret sauce to picking RAM other than being on this tested list, it's not widely known. Just finding a thread where someone could explain the problem and the fix was difficult.
I guess this is no different from any newly released thing. Better to wait until the quirks get worked out.
Having said that, getting ahold of the 3900x in the UK made me feel like I was too participating in a very limited market, situated somewhere in the middle of the Pacific ocean.
Motherboard is TUF B450M-PLUS gaming.
Chip is Ryzen 3600.
"My CPU still thought 0xFFFFFFFF was the randomest number ever, always, no matter what."
"What if two years later, I was still vulnerable to stack-smashing that I shouldn't have been, due to ASLR that wasn't actually randomizing?"
An excellent point... IF ASLR is dependent on RDRAND, and IF RDRAND always returns 0xFFFFFFFF, THEN ASLR will not work, and addresses will be laid out in memory deterministically rather than randomly... which is a big no-no for security...
A future OS would check the RNG's it uses for any security feature, on startup, and upon failure would stop, log the problem, advise the user, and ask them how they want to proceed...
The OS needs to paper over known defects, if consumers are to be kept sane.
I'm not saying you're wrong; I'm just saying, hey, there's "prior art" for this... <g>
long getRandomNumber() {
return 0xFFFFFFFF; // chosen by fair 4294967295 sided dice roll.
} // guaranteed to be random.There's a good related Dilbert too: https://dilbert.com/strip/2001-10-25
OpenBSD conveniently provides arc4random() in its libc for applications to use, and the same function is available for kernel components (obviously one needs to include different headers).
I know I'm nitpicking but I disagree with that statement.
There's a kernel option called CONFIG_RANDOM_TRUST_CPU which you can set to false. So if for some reason, even if it's a bad one, you don't want your random numbers generated by your cpu then that's that. End of discussion. (In theory, not sure if rdrand if trapable)
I get that you're skeptical about the quality of what's provided in /dev/(u)random because in most cases it's true. Should I ever feel the need to hookup a hw random generator then I hope programs would use that one instead of guessing they can do better by calling rdrand.
There's been a bit of drama around it:
On the other hand it is somewhat worrisome that brokenness of RDRAND can affect output of Linux's RNG that much.
There is assembly in WireGuard for the crypto primitive implementations, but those are generated by scripts and are based on either formally-proven implementations or highly-vetted ones.
[1]: https://elixir.bootlin.com/linux/v5.3.6/source/drivers/char/...
Other facilities in the kernel, such as ASLR, also use get_random_u32().
Many things in the kernel use get_random_u32(). That's the proper function to use.
When presented with this bug, the upstream kernel maintainers chose not to fix get_random_u32(), due to the availability (?) of microcode updates for AMD chips. That's not my decision. WireGuard is just a mere consumer of get_random_u32(), like all other modules. This is an upstream kernel bug.
So says a maintainer of WireGuard. HN is beautiful sometimes.
If RANDOM_TRUST_CPU is disabled, that will stop the kernel function from using RDRAND and avoid this issue for anyone using the ‘get_random_u32()’ function?
I just got my 3900x a couple week ago and I don't have this bug according to the test code provided in the article.
I'm running:
> Archlinux linux-5.3.7
> Linux-firmware: 20191022.2b016af-1
> Microcode: v2.2 (patch_level=0x08701013)
> MB: ASUS PRIME x570-P. bios=1201
I also find it kind of funny that he calls Asrock Asus every single time he mentions them in the article. If he was trying to install an Asus bios on his Asrock mobo then he's got bigger problems than this bug.
EDIT: looks like it's out of beta now - awesome! Thanks.
I'm pretty surprised that so much software seems to rely on external (be it from the CPU, or the OS) random number generation directly instead of simply using a software PRNG and seeding it with external sources (and maybe additional creative sources if desired). That seems like a way more predictable system to me, in the sense that at least you get guaranteed uniformly distributed numbers, even if the external seeds are not uniformly distributed at all.
Also, it is very easy to shoot yourself in the foot when being "creative" with random numbers, especially in the domain of security.
My conclusion would be exactly the opposite: If your RNG is so important that it has to be cryptographically secure, you owe it to your users to put in the time and effort of maintaining a proper implementation yourself, or at the very least use an open source library that provides this functionality in software. Otherwise you're always going to be at the mercy of a potentially misbehaving environment.
In terms of entropy, you don't really need to "maintain a pool" for CSPRNGs. You either have enough entropy to feed it with, or you don't. Once it is properly seeded, you can squeeze as many random bits out of it as you want (or at least, as many as anyone would ever reasonably need). It's really no different from a stream cipher, the key is the seed, and you're just encrypting zeroes. You don't need to suddenly get another randomly generated key after encrypting 100 MiB to encrypt the next 100 MiB securely.
Another great thing about entropy is that you can't reduce it. Which is why you really don't have to spend any time thinking about whether a particular entropy source is well behaved or uniformly distributed or anything like that at all. You just have to be certain that you have overall enough entropy that nobody can guess the entire seed. So anything the OS can give you? Dump it in there. Any kind of user interaction? Dump it in there. The time? CPU jitter? Network jitter? Just put it all in there. 100 MiBs of 0s? You know what why not, just put it on top because you literally can't make it worse, only better.
For example, seeding ChaCha with 256bits will give you 1 ZiB of output before cycling. That should keep you going for awhile.
A lot of PRNG are now implemented as the output of stream ciphers or block cipher in counter mode:
* https://en.wikipedia.org/wiki/Fortuna_(PRNG)
So 128 bits is all that is needed to get going.
Re-key every so often to ensure forward security in case there is a kernel-level compromise.
With AES-NI instructions in most CPUs, several GB/s can be achieved.
Wireguard’s mistake was not using what the OS gave them; I’m not sure why they didn’t.
If you have a secure stream cipher, you have a secure RNG. Not "can build", they're the same thing. If you only have block ciphers, the security proof is less trivial, but in practice "decrypt /dev/zero" still works.
If you don't have a secure (symmetric) cipher, you have bigger crypographic problems than bad random numbers.
And again, like before, the reaction to this is just hilarious to me. "It's a kernel bug." - Yeah, sure, but your users don't care, and your users shouldn't have to care. Because this an easy problem to fix, and if you're worth your salt as a developer you should fix it. Evidently Wireguard isn't worth it's salt.
This is completely true, but there is a property that your random numbers have to have that your seed doesn't need in any way: Being uniformly distributed.
If you need a random 64 bit number, you want every single bit in that number to have a 50% chance of being 0 or 1. This requirement doesn't exist at all for seed data, the entropy of all your seed data together just needs to be large enough for it to be computationally infeasible to calculate it from the outside.
I would say that is a pretty big difference. Also, if you just need uniformly distributed random IDs but those IDs aren't secret and don't need to be unpredictable, you don't even need a seed. Or at least, you can just use a constant as a seed. No interaction with the outside world needed at all.
Gee what if instead when it hit a threshold, it disabled the feature and called BUG()?
Other facilities in the kernel, such as ASLR, also use get_random_u32().
Many things in the kernel use get_random_u32(). That's the proper function to use.
When presented with this bug, the upstream kernel maintainers chose not to fix get_random_u32(), due to the availability (?) of microcode updates for AMD chips. That's not my decision. WireGuard is just a mere consumer of get_random_u32(), like all other modules. This is an upstream kernel bug.
Edit: There’s some interesting discussion in the Wikipedia article’s talk page’s criticism section.
We already know that parties like the NSA have intentionally backdoored RNGs that made it into production use.
If those supervisor cores want, they can literally just tamper with the entropy in your pool directly, they don't need to go through the charade of feeding it tampered entropy and then trying to un-mix it later.
It is far less risky detection wise to just tamper with the RNG directly.
Go ahead and try to tamper with that pool. On all existing and future kernel versions. After undoing kASLR.
Not arguing, that it is impossible, but malware writers usually choose easier paths. When Samsung backdoored their phones [1], they didn't make a pure TrustZone rootkit. Instead they created a companion app, acting as helper for performing high-level tasks within OS bounds. Intel ME rootkit has it's own network stack and can use Intel network card, because it is both easier and safer than trying to interact with constantly updating OS, written by different people.
A rootkit, trying to perform any kind of complex interactions with host OS, may be exploited and taken advantage of, — for example see bugs [2] in AMD's PSP, that allowed host OS to take over by giving it specially crafted certificate.
[1]: https://www.zdnet.com/article/backdoor-in-samsung-galaxy-dev...
I think this is the real kicker, and might represent one of the next major fronts in the security struggle. It's a little different from the debates happening right now about support periods, that at least has clear economic implications. It's one thing to argue about whether a product should still be supported at all. But it's quite another when something is being supported, and does in fact have a patch available, yet many owners still can not apply it anyway. That seems like an avoidable failure, and something worth considering legislation around. The industry could and should have more standardized methods and requirements to make sure that any patches that are created do make it out to product owners quickly and universally, there just hasn't been consistent motivation.
At least on Linux, that's not true. Intel publishes microcode updates on Github [1] (and distros package it) and AMD has it upstreamed in linux-firmware [2], so you don't have to rely on motherboard vendors at all.
[1] https://github.com/intel/Intel-Linux-Processor-Microcode-Dat...
[2] https://git.kernel.org/pub/scm/linux/kernel/git/firmware/lin...
No user space program should need to use rdrand directly at all.
That doesn't explain why systemd uses it, of course.
[1]: https://elixir.bootlin.com/linux/v5.3.6/source/drivers/char/...
In comparison get_random_u32() is safe to call at any point — including early boot — and does not affect global entropy pool. At worst it may return low-quality numbers, but that can be easily fixed by running your own peudo-random generator on top of it (which is a good idea anyway because you don't want your kernel module to contend with other parties for RNG ownership).
Indeed, because RDSEED is actually what most people want anyway.
Its an assembly instruction that gets the job done. People should be expecting that the assembly instructions of their CPUs work as intended. No different than using AVX-intrinsics or hand-crafted assembly in x264 / x265 code.
In any case, RDSEED is the assembly instruction for gathering entropy (aka: setting a random number generator should use RDSEED), while RDRAND is an older assembly instruction for purely getting a cryptographic random number. Its slightly different amounts of entropy involved in RDSEED vs RDRAND. So this is a very subtle issue that requires a lot of understanding of the x86 assembly instruction set.
But if you understand these details, then by golly you should use the instructions!
I'm sorry, but I really don't feel bad for this person.
Also, it's disingenuous to say that only "some" (the author uses that word) motherboards have updated microcode. "Most" would be more accurate, and "pretty much all, with few exceptions" is actually closer to the truth. Even my el cheapo A320 chipset motherboard already has AGESA 1.0.0.3 ABBA.
> I'm sorry, but I really don't feel bad for this person.
And this is why Linux will never take off as a consumer OS.
https://spectrum.ieee.org/computing/hardware/behind-intels-n...
https://www.hotchips.org/wp-content/uploads/hc_archives/hc23...
NB: RDSEED was added at a later date, some people didn't like that Intel used AES for conditioning, they wanted the raw bits.
sudo hexdump -C -n 64 /dev/urandom
sudo won't be required to read from /dev/urandom, although it probably is for /dev/hwrng.I just hope people don't think that every system build will deal with these problems.
It was a very frustrating thing to happen, since I was super excited to be back on AMD for my laptop.
Other facilities in the kernel, such as ASLR, also use get_random_u32().
Many things in the kernel use get_random_u32(). That's the proper function to use.
When presented with this bug, the upstream kernel maintainers chose not to fix get_random_u32(), due to the availability (?) of microcode updates for AMD chips. That's not my decision. WireGuard is just a mere consumer of get_random_u32(), like all other modules. This is an upstream kernel bug.
https://en.wikipedia.org/wiki/Shot_noise
Johnson Nyquist noise is measured from resistors, and is therefore present in every circuit. Shot noise however, is more evident in transistors.
------------
I don't know if the AMD circuit is Shot-noise or Johnson Nyquist noise. But its important to remember that there are many sources of noise / randomness in normal circuits.
The funny thing about true-random number generators: noise sources are all over the place! A beginner who plays with transistors will almost immediately "discover" a form of noise, and be forced to minimize it.
With regard to "true random" circuits (of which there are many, many different kinds), they all isolate a particular noise, and then amplify it. You have to isolate a particular form of noise if you want to get a good entropy estimation.
EDIT: It seems both shot noise and Johnson-Nyquist noise are white-noise, so maybe it doesn't matter. White noise + white noise should result in white noise, but its been a long time since I've taken this circuit class...
Which will affect the voltage and current on the line, which will affect the timing of the circuits, which will be amplified into uniform noise through some mechanism, generating true randomness. After all, we can't predict the location and velocity of molecules! Heisenberg uncertainty principle!
Then why are half the posts on Hacker News about how wonderful Ryzen is and how Intel is all but dead, etc.?
Also, this issue was discussed here when it was noticed shortly after release. AMD acknowledged the issue and released a fix pretty quickly, subject to microcode updates don't necessarily make it out to all motherboards very rapidly.
Since when? On Linux anyway microcode can be installed just fine as a package, often bundled in kernel updates.
( see the alt text on https://xkcd.com/221/ )
Each pixel will have a bias, but x-oring lots of them together gets nice, unbiased random bits to stir into your pool.
If it doesn't have a CCD, maybe it has a microphone or ambient-light sensor. That yields fewer total bits, but usually enough.
Maybe it has two, or all three. The advantage of mixing randomness from multiple sources is that you don't need to trust them all. If one starts feeding you FFFF..., the output has just a little less entropy, not none.
My guess for why AMD produces FFFFF... is that a Spook Mode was activated by accident. If so, it was supposed to switch to that only on command from the "management engine", the wired-in exploit every big CPU has. Maybe somebody deep within AMD wanted to ensure that we would not trust our RDRAND instruction overmuch.