ZenHammer: Rowhammer attacks on AMD Zen-based platforms
comsec.ethz.ch
comsec.ethz.ch
Now, if you are a cloud provider that provides VMs on multitenant hosts, your threat model may be different.
Either way, avoid machines without ECC. TRR was a lame duck even when Rowhammer was still fresh, and bits flipping in DRAM will not go away unless the economics in DRAM manufacturing change (e.g. not).
This can be verified easily on the AMD site by reading the CPU specifications.
For laptop CPUs, there has been a time interval between the beginning of 2022 and the autumn of 2023 when ECC support was specified for all mobile Rembrandt and Phoenix CPUs, but then the ECC support has been removed suddenly from the specifications of all non-Pro laptop CPUs.
- Desktop Ryzen CPUs support ECC, but implementation by motherboard vendors is not mandatory
- Laptop and G-series Ryzen CPUs only support ECC in pro variant
- Threadripper has ECC support
edit: not confirmed but supposedly laptop and APUs starting with 6000 series all support ECC.
For the current Ryzen 8000 laptop series (Hawk Point), the ECC support has been specified as missing from the beginning.
Last I checked ECC was not certified to work on most "consumer" oriented hardware, but AMD didn't make any attempt to actually disable it.
At least since Zen 3 (2020), all AMD desktop Ryzen CPUs have ECC support clearly included in their specifications, so it is official support, not just a non-disabled ECC.
This change was around the same time when Intel has begun to support ECC in some Alder Lake desktop CPUs (and in their successors), so it might have been a response to Intel's decision.
So now the ECC support depends strictly on the motherboard manufacturer. The best chance to find motherboards with ECC support is at ASUS and at ASRock (including ASRock Rack, which offers server boards for Ryzen CPUs).
https://www.asrock.com/mb/AMD/X670E%20Taichi/index.asp#Speci...
Supports DDR5 ECC/non-ECC, un-buffered memory up to 7800+(OC)The Riptide motherboard I use also has ECC listed now. https://pg.asrock.com/mb/AMD/B650E%20PG%20Riptide%20WiFi
Nevertheless that is still excessively high. While in the beginning for DDR5 there were only 80-bit modules, which could claim a +25% higher price, now there are 72-bit modules, like in the previous generations, which can justify at most a +12.5% price increase.
Trying again, I can find some Kingston sticks that are $120 each, so that's about 50% higher. Amazon's search is really bad, by the way. But that's not the "at most" price. And a month ago they were $140 each.
Also the spec sheet says they are 72 bit modules made out of 20 2GB x8 chips? That is baffling. Is 10% of the memory going unused? https://www.kingston.com/datasheets/KSM56E46BD8KM-32HA.pdf
Edit: Micron data sheets suggest that UDIMMs have 2x13 command/address pins and RDIMMs have 2x6, so that's one piece of the puzzle. Apparently UDIMMs can do x64 and x72, and RDIMMs can do x72 and x80.
Because the DDR5 channels must have a width of 32, 36 or 40 bits, with x8 chips one had to use 40-bit channels, even if only 36-bit channels are needed, so indeed 10% of the memory capacity remained unused.
Meanwhile, about a year ago, at least Micron has also introduced x4 chips. There are such ECC UDIMMs, using both x8 and x4 chips, which waste no memory.
On the market there are both modules made only with x8 chips, which do not use a part of the memory, and modules with a combination of x8 and x4 chips, without unused memory.
Asus motherboards for those CPUs also support it, as stated in the BIOS manuals I looked over. It requires changing the BIOS ECC setting from Auto to Enable.
I have done this on one such system, and the appropriate EDAC messages showed up in the Linux boot log.
That's not anything new though.
My local supplier (https://www.scorptec.com.au) has a fair amount of both, with ECC currently being about double the price (ugh).
When the AM4 generation was current, the difference in price was a lot less though. :/
Still it's worth it for piece of mind. Especially if you're undervolting the cpu. ;)
Overclock it yourself. It's not that hard.
Also I tried to get factory-OC'd RAM, but couldn't find any 32G sticks with ECC.
How much of that is it being actually slower and how much is it just ECC memory not being sold at pre-overclocked speeds?
This seems like all the argument necessary to require all computers - and I do mean all computers - to require ECC memory. The security risk is simply too great, and everything is too integrated not to make this change. Even a "gamer" on a pure gaming computer will have some crucial information on that machine, so I simply do not see how we've gone this far without making this change.
Which belongs to the gamer user, so it can be extracted or encrypted by every malware running on that system. No need for esoteric attacks like rowhammer at this point. Unless you think about, for example, exploring user machine via rowhammer in js running in a browser tab, but as far as I know that was never made practical.
* Are modern patrol read engines guided by the memory access patterns to respond to RowHammer style attacks?
* How aggressive would a patrol read engine have to scan the DRAM to safely stay ahead of RowHammer induced bit-flips?
* Would larger ECC words than the traditional 64+8 with multi-bit error correction change the game and allow us to build more reliable systems from DRAM with pattern vulnerabilities?
This is incredibly misleading. The paper they cite states:
When the ECC detection is used correctly 0.65%-7.42% of all bit flips still cause silent corruptions... On setup AMD-1, uncorrectable errors crash the system.
The attacker will need to cause dozens of machine halts in order to achieve even a single exploitable bitflip. Dozens of machine halts is not something that goes undetected.
Kudos for calling out JEDEC's terrible behavior on the rowhammer question, but we should not be downplaying ECC as a near-term solution.
Among consumer products, some AMD desktop CPUs and motherboards support ECC memory, and that's about it.
It's specifically mentioned on the ASRock motherboard pages under "Specifications". Some random examples:
• https://www.asrock.com/mb/AMD/B650%20Pro%20RS/index.asp#Spec...
• https://pg.asrock.com/mb/AMD/B650%20PG%20Lightning%20WiFi/in...
• https://www.asrock.com/mb/AMD/X670E%20Taichi/index.asp#Speci...
These all have:
Supports DDR5 ECC/non-ECC, un-buffered memory up to 7200+(OC)As a data point, I'm using a previous generation ASRock AM4 motherboard with ECC and that definitely works.
I'm undervolting my cpu and ram, and very occasionally (every 6 months or so?) one of those seems to be generating a correctable ECC error that gets propagated to warning messages on my terminal. Haven't bothered investigating any further though. ;)
For desktops it is much easier to choose ECC memory, because the additional cost (the cost of the memory modules is 50% higher for DDR5-4800) remains a small fraction of the cost of an entire computer.
What is needed is to buy a motherboard with ECC support.
An example of a good motherboard with ECC support is ASUS PRIME X670E-PRO WIFI (for AMD Ryzen). I have been using a similar ASUS motherboard with ECC memory from the previous X570 generation for the last 5 years and it still works fine.
There are several other such MBs, mainly at ASUS and ASRock.
For Intel Raptor Lake there are fewer and more expensive such motherboards, but they can be found at ASUS (Pro WS W680M-ACE SE) and at Supermicro, as "workstation motherboards".
But regardless, ECC still sounds the alarm when it's being attacked. If no one listens, there's not much ECC can do about that.
Only for some old modules, e.g. 5-years old or older, the frequency of errors can increase a lot, even up to many bit flips per day, which means that the offending module must be replaced.
This feature of identifying the aged modules is one of the main benefits of ECC.
I have not looked again at the AMD EDAC driver, which has been updated during last year, but previously, a couple of years ago, its feature of injecting errors for testing was broken on Ryzen (because it had not been updated since Bulldozer, at that time), so the only method to verify that error reporting is working was to overclock the memory in the BIOS settings, to ensure that errors will be generated. Obviously, for the test one should boot from read-only media, to avoid the corruption of the storage in the case of excessive errors.
If I knew how, I'd dial down the voltage very slowly while running a rowhammer PoC to either catch it hammering or catch edac counts.
> While the X570D4U-2L2T and its predecessors the X470D4U series supports ECC memory, there is a bit of a gotcha. As readers noted in the original X470D4U reviews, while ECC memory is supported and performing error correction, the reporting of that error correction was not functioning. In other words, even if you were experiencing continuous memory errors, no log of those errors was being recorded in the IPMI event log where one might expect them to show up. A user over on the Level One Techs forums had a conversation thread with someone from ASRock Rack, who reported that while the AM4 platform had ECC support, it did not have error reporting support.
[2] says:
> However we got AMD official respond today. AM4 does not support ECC error reporting function
[1]: https://www.servethehome.com/asrock-rack-x570d4u-2l2t-review...
[2]: https://forum.level1techs.com/t/asrock-rack-x470d4u2-2t/1475...
Standard ECC events are handled by the OS, and don't depend on (or otherwise interact with) an out-of-band management controller or any other external device. This works fine on Ryzen.
If you're targeting a specific machine, if you're throwing the exploit at a few thousand machines shotgun style then you're still going to get your botnet - it'll just be smaller.
Rowhammer and speculative execution attacks are incredibly labor-intensive and target-specific. They are targeted attacks for high-value targets.
Is there a process for the operations team managing the system to figure out that it was an attack and not just flaky hardware?
Normally a memory error does not happen more than a few times per year, unless you have a huge amount of memory.
Therefore when 2 memory correctable or uncorrectable errors happen in the same day, that should be enough to trigger an immediate report to the user or administrator of the computer that either there is an ongoing RowHammer attack that must be stopped or one of the memory modules is approaching its end-of-life due to aging and it must be replaced before it will begin to have very frequent memory errors.
At least on server computers it should be easy to configure their logging system so that a second memory error per day, even if it was correctable, should immediately send an e-mail message and/or an SMS to the administrator.
My understanding with spectre and meltdown was that it was an issue for escaping VMs and similar attacks - something AWS engineers should care about, but not me
This poses a significant risk as DRAM devices in the wild cannot easily be fixed, and previous work showed that Rowhammer attacks are practical, for example, in the browser, on smartphones, across VMs, and even over the network.Sure, it isn’t perfectly safe. If HN or my employer goes evil, they can rowhammer me I guess. I’d expect it to cause a big todo, though, so I’m not that worried about it.
I don’t really understand why people seem to think disabling JS is a big hassle. Is this motivated reasoning by web devs or something?
It is not a big problem, and the sort of “ambient shittiness” of the internet greatly improved by doing it. Most sites work fine, they’ll default to some (better) less dynamic state, maybe some ads won’t load. For those sites that don’t work, you can make an exception or leave. Personally I’m now mostly visiting sites by people who don’t enjoy over complicating things, and who think about fallbacks. It is great!
Your phone is selling your location for antiabortion fanatics to harass you, or help your stalker find you. Your ISP is selling your browsing history to anyone with a dollar.
That databroker that everyone was selling too just went to bankrupt and the banks are selling your data to anyone with a penny.
We desperately need wholesale privacy regulation.
Yet no privacy.
It's almost like the people who run the country,
do not believe in democracy.
The fact that I must run JavaScript written by just about anyone, in order to live in a modern society, or the fact that I keep having to write code in JavaScript in order to run a (completely non-JS related) business.
It takes only one creative genious to turn the next security issue into a thing that does affect us all. Some worm that eats all linuxes, a virus that spreads through all bsds or something that installs crypto miners on every second android or so. We cannot know.
And so we cannot defend ourselves against that. And so it's useless to worry about it. But it will happen. Our systems are way too monoculture, both soft- and hardware, to be protected against a digital potato famine.
mitigations=off
Other than that… we don’t often run random programs from the internet, right?
They’ve only scratched the surface for these sorts of bugs. Modern hardware is too complex to actually believe they’ll ever get them all.
Noscript would be much better of course, I guess I'm just too lazy to go that extra step.
At least, I typically see things about the trade off between usability and security and the need to enable certain use-cases. I think most security experts work in industry where their job is to figure out what can be done to patch things up within the constraint that their job doesn’t exist unless the company can do the stuff it needs to do to stay in business.
I don't really care about any of that, I just want to be able to read text from the internet without my system getting messed up. It is a much easier use-case, because static content is usually pretty safe (although I do think there have been vulnerabilities in font and image rendering libraries). We don’t need an expert to intelligently analyze things and balance against the interests of competing parties because there’s no need to push in the “open things up” direction for the most part.
Note that the average person wouldn't know WTF "DRAM" means, let alone "Rowhammer" or "Zen" or other esoteric industry terms.
Some of these exploits have been used in targeted attacks towards end users so the risk is not 0.
I've heard bitflipping a million times and never really got it (not that I made serious effort) until this.
https://comsec.ethz.ch/wp-content/files/hammertime_raid18.pd...
I feel like I just went through a 101 EE course. I had NO idea any of this was related to the actual hardware manufacturing imperfections, etc.
That explains the name Rowhammer. I've probably been under a rock and everyones knows this stuff.
> Due to the extreme density of modern DRAM arrays, small manufacturing imperfections can cause weak electrical coupling between neighboring cells. This, combined with the minuscule capacitance of such cells, means that every time a DRAM row is read from a bank, the memory cells in adjacent rows leak a small amount of charge. If this happens frequently enough between two refresh cycles, the affected cells can leak enough charge that their stored bit value will “flip”, a phenomenon known as “disturbance error” or more recently as Rowhammer.
This makes it sound like it's unavoidable and inherent to making DRAM. It isn't.
DRAM manufacturers have been pushing the limits to an extreme. That's why. Pursuit of profit. This is no different from Ford deciding the cost of settling Pinto lawsuits (from injuries and deaths) was less than the cost of fixing the car's design.
The smaller the charge, and the closer together the charges, the easier rowhammer attacks are. Also, the smaller and closer together the charges, the faster, cheaper, denser, and efficient your RAM gets.
There are mitigations, but they are pushed to the limit.
As others said, this isn't just about profits. It's about being able to compete on costs (i.e. being able to survive at all) and to compete on the best performance. This places the problem less at singular manufacturers and more at the whole industry.
So, with memory encryption you are safer.
So I guess DDR5 still has a little bit of time here. Anyone know if this also affects LPDDR5x?
> Furthermore, we show that ZenHammer can trigger Rowhammer bit flips on a DDR5 device for the first time.
Was its design pushing material sciences such that the theory worked, but practical implementation required the 'crutch' of ECC?
Linus was right about ECC being needed, with higher capacities and speeds and reduced feature size it's becoming a must.