Try to make sudo less vulnerable to Rowhammer attacks
github.com
github.com
Do not want.
Rowhammer is a hardware problem ---- defective RAM --- not a software one.
The sooner everyone starts returning defective RAM and putting pressure on the hardware manufacturers to maintain correctness, the sooner we can stop this descent into insanity.
"They can always work around it in software" is the attitude that let Rowhammer exist, and continuing to fulfill that expectation will only make things worse.
It always amazes me how people can be so confident yet so wrong.
It's a problem of physics - there's various ways to try to mitigate it but the only way to completely avoid it would probably be to use SRAM and that is going to be extremely expensive when talking 16GB and not nearly dense enough.
It's not some conspiracy by "Big RAM"
Why wouldn't storing a cryptographically-secure checksum on every RAM row work?
I mean, if you had a DRAM labeled "DDR4-3200" but it can only work at a much lower speed (say DDR4-2400), it's clearly a problem of physics - the gate capacitance is too high, the driver transistors are not strong enough. And yet my reaction would be to take that RAM back to store and get my money back, not to defend manufacturers which claim false things about their chips.
And yet, when the bit flip problem caused by physics got so bad in DDR5 it couldn't be ignored they did fix it - by adding error correction codes. It wasn't that expensive to do so. Notice that HDD's hit the same problem as the got denser, and solved it in the same way (ie, by throwing lots of ECC at it).
I agree with the original poster. It's a hardware problem, caused by the manufacturers pushing the limits. And it's their problem to fix, which they can do. DRAM that doesn't corrupt itself isn't a big ask.
Yes it's a problem of physics and it's because they are trying to make DRAM too dense.
It's a problem of physics - there's various ways to try to mitigate it but the only way to completely avoid it would probably be to use SRAM
It's not some conspiracy by "Big RAM"
Look at the evidence. This didn't start showing up until around 2009-2010, and the industry managed to convince authors of widely-used memory testing tools to downplay the severity and/or not enable RH tests by default, because they didn't want the truth to be known that almost all RAM is defective. It might not be a conspiracy, but it sure is corporate greed.
Would you rather pay a little more for RAM that will work correctly under all access patterns, or RAM that is certain to produce bit errors under some conditions that can be encountered in practice? Unfortunately, with newer DDR3 and later generations, it seems you don't get a choice.
The choice just boils down to "buy pricey server boards/chipsets that have ECC RAM available" or "get bent".
ECC still superior by a longshot, this information notwithstanding.
I recommend [1] as an introduction to the semiconductor physics behind the Rowhammer problem. Rowhammer is an instance of the "weird machine" problem behind many security problems, i.e. a mismatch between two abstractions: the abstraction we pretend describes the system, vs the reality of the system. In the case of Rowhammer, that is the abstraction of memory as a digital device, against the reality of storing bits with capacitors and wires, ie. analog devices. Clearly a leaky abstraction. The denser you pack those capacitors and wires, the more leaky.
[1] A. J. Walker, S. Lee, D. Beery, On DRAM Rowhammer and the Physics of Insecurity. https://ieeexplore.ieee.org/document/9366976
"Capacitor plague of 2000 was a mismatch between two abstractions: the abstraction that capacitor actually provides datasheet-described amount of capacitance vs the reality of the system"
"Toyota unintended acceleration was a mismatch between two abstractions: the abstraction that ECU properly responds to accelerator pedal release vs the reality of the system"
Yes, digital systems are made of analog parts, but that's not a reason to accept systems behaving out of spec. For the last 50 years, the specifications for RAM have been pretty clear: as long as all datasheet requirements are obeyed, the only way to change stored data in one location should be to do a write to that location. If a memory chip does not act according to its own datasheet, it's not a "leaky abstraction", it's a hardware bug.
(Now, can this be fixed economically? I don't know, I could believe the answer is "no". However, the solution in this case is not software workarounds, but rather to make a new spec: "RH-RAM is like regular RAM but cannot tolerate certain access pattern")
https://blog.adacore.com/adacore-enhances-gcc-security-with-... and https://gcc.gnu.org/onlinedocs/gcc/Common-Type-Attributes.ht...
It's very relevant! The problematic comparison in this code isn't true/false! A feature that only protects true/false does not help here.
#define AUTH_SUCCESS 0x52a2925 /\* 0101001010100010100100100101 */
#define AUTH_FAILURE 0xad5d6da /* 1010110101011101011011011010 */
#define AUTH_INTR 0x69d61fc8 /* 1101001110101100001111111001000 */
#define AUTH_ERROR 0x1629e037 /* 0010110001010011110000000110111 */
#define AUTH_NONINTERACTIVE 0x1fc8d3ac /* 11111110010001101001110101100 \*/
going to see how i can work this into a project :) #define AUTH_SUCCESS 0x52a2925 /* 0101001010100010100100100101 */
#define AUTH_FAILURE 0xad5d6da /* 1010110101011101011011011010 */
#define AUTH_INTR 0x69d61fc8 /* 1101001110101100001111111001000 */
#define AUTH_ERROR 0x1629e037 /* 0010110001010011110000000110111 */
#define AUTH_NONINTERACTIVE 0x1fc8d3ac /* 11111110010001101001110101100 */
AUTH_FAILURE is still just !AUTH_SUCCESS (and almost a palindrome) #define AUTH_SUCCESS 0x052a2925 /* 00000101001010100010100100100101 */
#define AUTH_FAILURE 0x0ad5d6da /* 00001010110101011101011011011010 */
Doing a '!' operation in C on one of them won't yield the other value unless you also zero out the top 4 bits. Close enough, though... I still enjoy the symmetry as you did.Anyway, I'm curious why three of the values they chose have all zeroes for the top 4 bits. I wonder if there's a security-related reason for that.
Because rowhammer is attacking the physical memory structure, it can’t function at the level that knows what AUTH_SUCCESS is.
This attack just targets raw bits, so we need to protect these crucial state variables from bit-flips.
the distance between success and failure is 28
Very nice indeed. Such a simple mitigation and it makes evil people sad, which makes me happy.
If the attacker is already running code on your system, you kind of lost anyway.
Thus you have to figure out how to otherwise handle an enum whereby you can be reasonably assured of its value even with a flipped bit, hence the hamming distance.
SUCC FAIL INTR ERR NONI
0 28 20 11 16 AUTH_SUCCESS
28 0 12 19 14 AUTH_FAILURE
20 12 0 31 16 AUTH_INTR
11 19 31 0 15 AUTH_ERROR
16 14 16 15 0 AUTH_NONINTERACTIVE
Sure, AUTH_SUCCESS and AUTH_FAILURE have a Hamming distance of 28, but it takes only 11 or 16 bit flips to go from AUTH_ERROR or AUTH_NONINTERACTIVE to AUTH_SUCCESS. (AUTH_ERROR can only happen from an internal error, so I believe AUTH_NONINTERACTIVE is easier to trigger.)A quick Python search was able to find some alternatives:
0x0f7b74c5 0x810d2b99 0x63a64616 0xcab4a865 0xbe705abb
...maximizes all distances (17--19)
0x28d803a4 0x352ef6d3 0xdb61dce1 0xb3edf85c 0xe62f7508
...maximizes a distance from the first and others (21--22), disregarding other pairs (14--21)
It seems that fixing one element to be a bitwise negation of the first element is not a good search tactic in my short testing. Also as notpushkin noted, if you really want to disregard other pairs you should just make one pair with the maximal distance and derive every other code from them (say, -1 0 1 2 3 would work for this purpose).By the way, finding a binary code with maximal Hamming distance is an open problem [1] [2].
[1] https://www.win.tue.nl/%7Eaeb/codes/binary-1.html
[2] https://math.stackexchange.com/questions/4288902/generation-...
This was my next question. It'd be great if there was an algorithm for finding N codes as close to equidistant as possible.
I've sometimes contemplated the possibility of doing things like this to guard against memory errors causing mis-entry to particularly critical control flow paths - this is certainly an example of that. But never heard of anyone actually trying to do this until now.
A "how to write rowhammer-resistant code" writeup would definitely be useful here - even if it is definitely something people cannot do for anything, I can certainly see cases where there is a case for it.
The feedback from the rust subreddit was basically protecting the bool but not the if statement is of limited use. Potentially it can even make it worse, since there is now more code being executed that might become bitflipped.
That inspired me to make a crate that periodically checksums your program code while it's running, to make sure it hasn't changed. Got it working on Windows and Linux, but then it ended up like most side projects. Maybe I should polish it up and publish it.
I was not able to test any of my Raspberry Pis because the test code used some facility available on the AMD64 architecture that is not available on the ARM64 processors. The newer Pi 4B and 5 use modules with on die ECC.
It seems to me that ECC should prevent Rowhammer susceptibility. That should prevent it on server grade H/W for anything still in service and newer consumer systems.
I have no idea if Rowhammer affects other architectures than AMD64.
The paper linked in the patch references other works showing every defence at rowhammer can be bypassed somehow (I've not followed them all) - e.g. it specifically says that ECC and the like can be bypassed.
I also did not expect that the various defenses, including ECC, could be bypassed.
The attacker cannot control precisely which bits will be erroneous. When much more than 2 bits become erroneous, in a small fraction of the cases no error will be detected but a wrong value will be read at the next access.
However, in the majority of the cases an error will be detected, either non-correctable, or correctable in which case the corrected value will be wrong.
Despite the fact that wrong corrections are possible, in a system with ECC that is configured correctly it should be impossible for a RowHammer attack to escape detection, unlike for a system without ECC memory.
On a computer that is not defective, memory errors happen very seldom, typically one error after many months. Even only 2 correctable errors that happen in the same day represent an event that can be explained only by either a RowHammer attack or by a memory module that has become defective.
Therefore, a well configured computer with ECC memory should alert immediately its administrator when 2 on more errors happen in the same day, even if they had been correctable errors, because this requires immediate action, either stopping a RowHammer attack or replacing the defective memory module.
It would be pretty much impossible for any RowHammer attack to attain its target without triggering 2 or more ECC errors, which will reveal the attack attempt.
Only when there is no ECC the attack can proceed undetected for a time long enough to be successful.
> with RowHammer protection mechanisms disabled
I wonder what this means. Is it S/W mitigations or does it include H/W factors like disabling on-die ECC.
It makes sense to me that with all other things being equal that higher density would lead to more susceptibility to Rowhammer. But as always, other things are not equal. I expect that on-die ECC would reduce susceptibility to Rowhammer and AFAIK that is used for DDR4 and DDR5 RAM, but perhaps not exclusively. Or did disabling "protection mechanisms" include disabling that (if it is even possible.)
I don't believe that Rowhammer mitigations happen inside the DRAM chips themselves, I think that they are being put into the memory controller that talks to DRAM. Since DRAMs with built-in Rowhammer defences would have to spend transistors on this defence, those transistors would be 'wasted' in situations where Rowhammer is not part of the attacker model.
The disadvantage is that the controller and memory are made by different companies, so standards are required to agree on what access patterns are acceptable.
That's not a particularly big assist, and doesn't sound hard to mitigate.
If TRR pushed the rows it refreshes 10% closer to triggering their own TRR, then those 300,000 accesses would have triggered multiple refreshes in the target row.
It's easy enough to deal with if they stop cheaping out. Each row is so wide. A 10 bit counter to trigger neighbor refreshes would barely take any space.
> gcc -DRND1=0x$(openssl rand -hex 4) ...
That would cause grief to reproducible distro initiatives.
It is perfectly good enough for the error code enumeration to be statically randomized into hard coded constants. The attacker is very unlikely to flip every single bit of one valid value so that it resembles another valid value.
Even if the values were randomized at compile time, if the executable is readable to the attacker, the attacker can learn what those values are.
If the executable is not readable to the attacker, the attacker can just pull a copy of the executable from the distro package: executables are installed from widely used binary packages, not freshly compiled for every system.
A comment points out that they aren't randomized:
> The values used were chosen such that it takes a large number of bit flips to change from allowed to denied. Using random values doesn't really protect against this attack.
Preventing known attacks is not pointless at all.
Just inefficient; a jump table optimization is impossible on the values. Speed is sacrificed for security. A jump table itself could be row-hammered to jump where the attacker wants!
I also believe those use-cases aren't that common anymore since multi-user systems fell out of favor. There is an argument that most of us could use a vastly simpler tool instead to reduce the attack surface. But that tool wouldn't be sudo, because sudo is built around supporting all these use cases.
And we have the OpenBSD folks focused on clarity and security.
Apologies for self promotion, but I wrote a relevant blog post that discusses this[0]. Is there any way of mitigating this trivial attack?
I feel like the Unix/Linux security model is broken.
This is such a crazy hack to get some protection around key variables - by requiring 32-bit manipulation.
Why isn’t this just “done” on the phy layer, or somehow detected automatically in-flight? Does the compiler protect against it? At what performance penalty?
It’s really much nicer just not thinking too much about it and going on believing our bits.
I wonder, is choosing random enum values at runtime more secure against Rowhammer than just having fixed values that were chosen randomly once and compiled in, since presumably the attacking code now has no way to know which bits it needs to flip? If so, it might even be desirable to implement this as a "secure enum" in a compiled language.
It would be neat to see an algorithm that generates suitable values.
Although I'm not sure if it's optimal, the many case seems to be the same as the 2-case but repeated for every 2 items. I expect they double checked that the amount of bitflips is still pretty high.
Maybe a better algorithm for the many case would be something like the popcnt parallel patterns:
* 0b0101010101010101
* 0b0011001100110011
* 0b0000111100001111
* 0b0000000011111111
Since they would all have equal hamming distance between each of the entries.
And indeed, the two values ate bitwise complements.
Making something like this into a panic is not a good fit for Rust as-is. Because enums are proven to have only correct values, not only is code written to assume pattern matches cannot panic, but compilers are free to optimize around only having valid values as well. That goes not only for the enum discriminant, but for any associated values being properly initialized values of their respective types.
In a sense, Rust lets you write code as if invalid values never happen, so there's less to check for in your code. It's understandable from the perspective of the abstraction needed for computer code to be "correct" and not just temporarily getting away with Undefined Behavior. There are simpler ways to violate it than just rowhammer, write straight to process memory for example, which can also violate invariants that compilers assumed while optimizing.
If you wanted to compile Rust (or anything else) with a hardening mode that does check what should be redundant values, it would be a lot slower and code that never panicked before would now panic, but it would probably be a worthwhile tradeoff for some programs to opt into. After all, if you built for CHERI or arm64e and got a machine exception from an unauthenticated pointer, you'd be thrilled you mitigated a vulnerability even if it violated your higher-level language model. Defense in depth and all that.
Maybe someone feels motivated enough to write an RFC and prototype for this. It just wouldn't stop at enum values, it should mean all sorts of other things too, such as not eliding any other checks that appear redundant given assumptions like immutability. That's what makes it slow and hard to reason about.
__attribute__((rand)) enum e { FOO, BAR, ... };
which randomizes the values, as an extension.You only need this in specific places, like setuid programs.
Randomization can be bad because it wrecks build reproducibility; it would have to be tied to the GNU Build ID.
If such an enum is used in any interface between files, the randomization has to be the same in every translation unit.
Maybe the syntax could specify a seed: rand(42).
Which is why only a couple of the values are random, with the others being those values, XOR'd with 0xff*
(void)strlcpy(des_pass, pass,sizeof(des_pass));
As a full-time Linux user, I haven't used sudo for years. Rather, I do 'su root'. However, I noticed several years ago (Debian) that upgrades would iterate twice, seemingly accommodating two accounts. Emulating a moron as I do, I never exerted the effort to learn why. I simply began, after 'su root', entering 'sudo su', which despite always having sudo disabled, seems to make me proper root.
I'll often use synaptic package manager when I want a cleaner, easier interface to explore packages. If I only 'su root' it won't open unless I append .... something similar to 'pkex' to the end, but if I do 'sudo su', I can run it using only 'synaptic'. Regardless, I refuse to use sudo otherwise, even when I'm emulating something sentient.
Be afraid. There are morons using Linux, and some are quite productive despite.
(The man page recommends using --login over the single dash, but it also says they are equivalent. Maybe I'm too much of a moron to understand the difference, but the single dash is less typing)
By maximising the Hamming distance between binary values used in enumerations (enum variables) and boolean variables, one could detect if the code has entered an invalid execution path and then trigger a watchdog reset, e.g. an else statement or a switch statement with a fall through case that would not normally be reachable.
Scroll to the top for the reference to the Mayhem attack it's trying to guard against.
Cons: humans (and some software) expecting "sudo" will work while interacting with your system.
It would seem to be particularly dangerous, but then I don't see everybody getting pwned.
How?
And we know hardware bugs are real, that's the whole point of rowhammer.
You might want to look at Facebook research:
> “Silent data corruptions are not limited to rare one in a million occurrences within a large-scale infrastructure. These errors are systemic and are not as well understood as the other failure modes like Machine Check Exceptions.” A large part of the responsibility should be shared by device makers, Facebook says.
https://www.nextplatform.com/2021/03/01/facebook-architects-...
But I use 7zip as a library, so if I get a CRC error I still have the output. Which in my case is JSON files. By careful diffing and going over them I could identify single bit flips like '{"base": 10}' being decompressed to '{"bbse": 10}'. I made this example up, letter 'a' becoming 'b' might be a multi-bit flip, but you get the idea.
https://en.wikipedia.org/wiki/Row_hammer
The initial research into the row hammer effect, published in June 2014, described the nature of disturbance errors and indicated the potential for constructing an attack, but did not provide any examples of a working security exploit. [1]
[1] (June 24, 2014). "Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors"By my recollection there was a discussion of rowhammer and making it work on a (original) Freenode channel circa 2010 (or earlier) in response to a related thread on a reddit security hacking subreddit.
ie: it was being discussed in public channels some four years prior to a paper cited as "initial research".
Addendum: Mind you, lots of things get kicked about and implemented before actual papers appear on them for the first time in public.
Mind you, lots of things get kicked about and implemented before actual papers appear on them for the first time in public.
I was thinking of many examples I know where a technique is developed and used in industry (mineral exploration | remote imaging | secret spook stuff) and much later (five years or more) gets a first mention in acedemia .. where they may or may not get a robust working version happening .. just a rickety proof of concept.As silicon density increased, the issue became more urgent.
I suspect "observed in fabrication lab | not disclosed" dates back some years before the paper .. once observed there's always a path to exploitation - but why would anyone broadcast that?
By the time it was chit chat on IRC the general feeling was that some TLA has a working exploit. (obviously unpublished).