Researchers have significantly increased the scope of the Rowhammer threat
wired.com
wired.com
The problem of corruption at the physical layer when certain types of bit pattern occur is also encountered when transmitting data over a wire; constraining the physical parameters to remain suitable for the naive representation of binary works well for getting a signal across the PCB, but would be extremely limiting at intranet scales. The usual approach is to modulate the data in a way that avoids encoding the problematic bit patterns.
Are the tradeoffs necessary to maintain the simple abstraction worth it in this case? I don't know, but considering how much of a bottleneck RAM has become for modern hardware, I think it's worth considering the alternatives.
From a different angle: I think your point is fair but I also think that for it to apply to this situation, the memory vendors would have needed to loudly and openly say that they were invoking that tradeoff so the OS vendors could adjust. Presumably that would also result in a lot of benchmarking being done to see if the net effect of a physical-layer vulnerability and a software-layer mitigation was actually a net positive.
If we can make volatile memory chips significantly more dense by letting them be lossy, then lets either add another layer to the memory hierarchy and/or rename L3 cache to RAM and move the L3<->L4 mechanics into real software.
At any rate, manufacturers shouldn't just be silently eroding the abstraction so they can compete on density harder.
There is nothing wrong with selling RAM where certain access patterns corrupt the content in predictable ways. There is everything wrong though with selling that RAM for use in systems that are known to expect RAM to return exactly the bits written to it with a certain (high) degree of reliability. And it is wrong precisely because it is not a tradeoff. If you are honest about the properties of the RAM you are selling, then that is the basis for the system designer to make a decision whether using your RAM with an appropriate interface is a better choice than using more reliable RAM with a "traditional interface". Pretending that your RAM is suitable for the "traditional interface" is what prevents the tradeoff from happening and is essentially fraudulent.
RAM is pretty expensive.
It seems the market doesn’t agree with you. Why do you think that is?
Also known as race to the bottom.
The serious impacts show up when you can't rely on it being done properly, and have to use expensive workarounds.
It's still a correctness issue today, too. I don't understand why manufacturers (and their customers) consider it OK to ship broken DRAM chips that do not conform to their stated specifications.
Rowhammer isn't (just) a security issue to be worked around, it's a hardware bug that needs to be fixed. As far as I can tell, it hasn't been.
Because they can, and sucks to be you. This is how things are everywhere. For competitive markets, the only real quality pressure is regulatory and contractual (and maybe reputational, sometimes). There needs to be a direct feedback loop between the value end-customers care about and the profit of producers/sellers for that value to matter.
As a random and interesting example of this phenomenon (really seen everywhere), here's something I learned yesterday: according to Derek Lowe[0], there's no graphene supplier anywhere that actually supplies you graphene, and they all tend to lie about it. Apparently this is one of the big things that holds graphene research back (and probably invalidates a bunch of papers).
--
[0] - http://blogs.sciencemag.org/pipeline/archives/2018/10/11/gra...
Pretty much all tech is this way. Layer 1 of most copper, fiber and RF networks and long buses require scrambling[1] of the data to prevent issues caused by clumps of 1s and 0s. Modern x64 CPUs scramble[2] data before its written to ram. SSDs scramble[3] data before writing it to the physical flash chips.
[1]: 8b10b and newer techniques
[2]: https://web.eecs.umich.edu/~misiker/resources/HPCA17-coldboo...
Even if the attacker was able to get the flipping completely reliable, there would presumably be a learning/probing phase with a period of elevated ECC. Either this probe could be detected, or the attacker would be forced to remain below a threshold of detectability slowing the attack down enough to make it impractical?
/sys/devices/system/edac/mc/mc0/ce_count
The problem with characterising it as an "attack" is that it leads to the notion that certain access patterns are "bad", and that's not a slippery slope we should be heading down...
Sure, that would reduce the total ram by 1/8 ... But that would be a design choice to implement. Is ECC ram only 12.5% more expensive than non-ECC? If its higher, it may indeed be more advantageous to use non-ECC -if- a software compensation can be implemented.
In either case, extra memory accesses would be needed since checksums need to be loaded from memory. This would also make cache misses more frequent, since checksum data would evict non-checksum data from cache constantly. This would have a huge performance impact - most software contains a LOT of memory accesses.
However, it might be feasible to mitigate this in specific cases by having custom code in software that needs to be secure.
These checksums are typically done on blocks of payload data of course, not all memory content.
Most don't :)
Although: I once investigated a soft freeze on a realtime-patched Linux system that turned out to be caused by a vendor's software somehow managing to indefinitely stall an RCU grace period, eventually consuming all available memory on the system. The kernel core dump being over 4GB in size was a bit of a give-away.
[1] https://www.cs.vu.nl/~herbertb/download/papers/throwhammer_a...
Also for people who just want the link of the academic article (including abstract):
https://cs.vu.nl/~lcr220/ecc/ecc-rh-paper-eccploit-press-pre...
However, it's my understanding that exploits depend on running code (including JavaScript) on the target system (or in a sandbox or VM). Is that true?
I haven't read the paper, so I don't know how reliably they can do it in a real world setting where they are not the only people interacting with the server, but they demonstrate that it's possible.
But isn't a key-value server perilously close to a database prompt? And this exploit depends on having authenticated access, right? Otherwise something like fail2ban would prevent hammering, I'd think.
It then seems somewhat plausible that such packet data could effect faulty RAM.
They mentioned the attack can work with roughly a week worth of unprivileged runtime, as long as the ECC mode of the ram chips in the targeted system has been previously sufficiently reverse engineered.
Is that too alarmist? To me, it sounds like something perhaps too cumbersome for casual drive by attacks, but it seems right down the alley of so called "persistent threats", or whatever it is we call those guys nowadays.