(Although linked in that thread systemd code has fallback anyway, so I’m not sure how it fails at all).
Edit: not to mention that it’s better just to not use it, ever. Quite sane and sensible thing.
(Although linked in that thread systemd code has fallback anyway, so I’m not sure how it fails at all).
Edit: not to mention that it’s better just to not use it, ever. Quite sane and sensible thing.
* https://github.com/systemd/systemd/issues/11810#issuecomment...
H. Peter Anvin's educated guess was that some MSR flag, that has the effect of controlling this instruction, is not being saved and restored in the processor across a suspend, and the result is a processor state where it signals that the instruction is succeeding but it is not actually returning random values.
Let’s wait for clarifications how that person has done the tests.
However: the bug is about systemd failing to get entropy, not getting nonsense entropy.
* https://github.com/systemd/systemd/issues/11810#issuecomment...
There is no "however". This bug is about code that is, according to the AMD doco, using the instruction correctly; but that is, because the AMD processor has this possible state after a suspend+resume, getting all-ones as its random data, thereby causing ID collisions in a fairly wide range of possible things from freshly re-generated machine IDs to journal file header block IDs, and including unit invocation IDs.
* https://github.com/systemd/systemd/blob/717e8eda77b93ac396dc...
* https://github.com/systemd/systemd/blob/717e8eda77b93ac396dc...
* https://github.com/systemd/systemd/blob/717e8eda77b93ac396dc...
This indicates that a "should be fine" in another comment is not in fact true. (-:
* https://github.com/systemd/systemd/blob/717e8eda77b93ac396dc...
Why would OpenSSL fail visibly if the API was returning success but with non-random data?
That would be quite something to read the comment on;
// Here we check to be sure the Earth is not flat
...
// Seeing that the Earth is indeed flat, we will iterate in this absurdity some more in the hopes it rounds out eventually
Not sure how the kernel devs generally go about testing the “this virtually never happens” code paths without adding debug switches to every unhappy path.
Certainly I doubt they are using DI/IoC to wrap an interface to RDRAND which allows unit testing the failure modes.
At least the result is a failure to generate a key, not a compromised key.
Chip vendors, and the tools they use to design them, actually do a significant amount of work for error cases. This is important not just for correctness, but for production yield, reliability, temperature and radiation hardness, etc. As chips get larger and more dense this becomes more and more important.
Single Event Upset (where one bit flips) is an example of the type of error. https://en.wikipedia.org/wiki/Single_event_upset
No; it turns out that's giving systemd too much credit (sadly). See [1].
The problem appears to be that RDRAND was signalling success, but producing a nonrandom value. This is bad and a violation of the specification.
Can't speak to Linux kernel development, and in this particular case, that isn't the problem.
The linked bug involves systemd using the world's worst random number generator. A security engineer goes into more detail on this twitter thread[1]: https://twitter.com/FiloSottile/status/1125840275346198529 (or unrolled: https://threadreaderapp.com/thread/1125840275346198529.html?... ).
> At least the result is a failure to generate a key, not a compromised key.
In fact, the result is a compromised key -- the bug report is due to colliding "globally unique" identifies generated through a flawed random gathering process.