Defending against Rowhammer in the Linux kernel
lwn.net
lwn.net
Lunus covers the TL;DR in this quote "there is nothing remotely sane you can do in software to actually fix this." and gets down to ECC memory being the golden solution.
Security is all about increasing those difficulty thresholds, not absolutes.
TRR might better fit the description of a "golden solution" although it is still technically a mitigation. I would very much like to see more ECC in consumer/desktop systems.
Now with these public issues becoming more open, one can only hope that some consumer product moves to ECC and ends up driving ECC memory prices down in much the same way as mobile phones have driven many sensors and associated hardware down in price, thru volume.
https://www.youtube.com/watch?v=dfIoKgw65I0
https://www.blackhat.com/docs/us-15/materials/us-15-Herath-T...
There are a lot of details to consider. E.g. reading DRAM (rather than from on-chip cache) consumes a lot more power, which could be quite harmful on a laptop.
It is tested on on an Intel SandyBridge CPU. For a full deployment, we would need to know which bits of the physical address are used to select the DRAM banks and rows for each specific CPU. There has been some effort to reverse engineer these mappings by various people. Two excellent sources regarding these mappings: http://lackingrhoticity.blogspot.com/2015/05/how-physical-ad... https://www.usenix.org/system/files/conference/usenixsecurit...
In any event, the kernel can know what's behind the curtain and that is the context in which the suggestion and news item exist.
It //STILL// sounds like something that happens in the MMU, at the time that the virtual address is mapped back to physical addresses via the page tables; which means that the kernel still knows the real backing addresses.
Here are more details of how it actually happens: http://lackingrhoticity.blogspot.ca/2015/05/how-physical-add...
dmidecode seems to indicate that the description from the actual RAM is lacking transparency about it's internal geometry.
I can get the ranking for slot-level interleaving from dmidecode on my systems (which means a kernel could or already has it as well).
Thinking about the inside-chip geometry issue as well as the current on-list proposal in the news item I've reached a different potential solution.
If page faults are tracked /by process/, the top N faulting processes could be put in to a mode where the following happens on a page fault:
* Some semi-random process (theoretically not discover-able or predictable by user processes) picks a number between 1 and some limit (say 16).
* The faulting page, and the next N pages, are read in (which should pull them in to the cache).
This would help by making it harder to mount a /successive/ attack on different memory locations. Processes chewing through large arrays of objects in memory legitimately shouldn't be impacted that much by such a solution; surely less so than a hard pause on the process.
Am I missing some aspect of page mapping/cache management?
https://www.usenix.org/conference/usenixsecurity16/technical...
The memory controller incorporates a DDR3 Data Scrambling
feature to minimize the impact of excessive di/dt on
the platform DDR3 VRs due to successive 1s and 0s on
the data bus. Past experience has demonstrated that
traffic on the data bus is not random and can have
energy concentrated at specific spectral harmonics
creating high di/dt that is generally limited by data
patterns that excite resonance between the package
inductance and on-die capacitances. As a result, the
memory controller uses a data scrambling feature to
create pseudo-random patterns on the DDR3 data bus to
reduce the impact of any excessive di/dt.
So basically, all the kernel knows is that the data is in there "somewhere", and can be accessed traditionally.[1] http://www.intel.com/content/dam/www/public/us/en/documents/...
I designed a number of DRAM memory boards "back in the day", but haven't kept up with recent developments. But this idea could be used as a starting point by someone more in tune with current hardware to write a kernel module to help mitigate Rowhammer.
One key thing to know is that, internally, a DRAM chip isn't accessed by a single row (of let's say 32 bits) at a time. What happens is that a read causes a large number of bits (literally thousands) to be accessed and refreshed at once. Then the selected 32 bits are returned to the CPU. But, as a side effect, all those 1024 bits (or perhaps a lot more in current DRAM chips) are refreshed.
So what's needed is a background process that does the following for all of physical memory:
perform an uncacheable read of 32 bits direct from DRAM
increment read address by perhaps 32 words
pause for some small amount of time (perhaps 1 usec)
repeat forever
This task of repeatedly sweeping through physical memory will, as a side effect, cause all memory cells to be refreshed.Obviously there is some magic needed, which can only be done in the kernel. First, all of physical memory must be able to be accessed by this process. Second, some tuning must be done to keep the task from consuming too much memory bandwidth. Third, it might make more sense to do something like reading quickly a burst from 4 different physical memory locations then pausing for 4x as long.
Unfortunately, running this type of program would be devastating in terms of power consumption. DRAM chips consume much more power while being accessed than while they are in standby. So it probably would have an large deleterious effect on a laptop. But it would probably be OK on a desktop or server.
That just the basic idea. There is a lot of tuning that could be done. For example, instead of reading thru all of physical memory, perhaps just read only the kernel memory. That's a lot less memory, a lot lower power consumption. The idea is that corrupted kernel memory is potentially a lot more harmful than corrupted memory used by some random user process.
"For example, the current generation of chips (DDR SDRAM) has a refresh time of 64 ms and 8,192 rows, so the refresh cycle interval is 7.8 μs."
Sounds like your idea is already a thing! That's cool! Unfortunately it seems like the performance-cycle period tradeoff is difficult and Rowhammer is taking advantage of that.
Protecting subsets of memory is an interesting idea but what bits do you choose? What if I could Rowhammer the memory frame behind a COW page belonging to a setuid binary? I feel like it might be an all or nothing kind of thing in this situation. Who knows though, maybe there's merit in specifically protecting kernel owned frames, it's difficult to say for sure.
For ECC memory systems your idea is called ECC scrub. The idea is to trigger SBE correction before more bit flips occur, turning it into an unrecoverable or undetectable error. Usually it appears in BIOS as Patrol Scrub (continuously and slowly walk through all memory, slow enough not to contend with active programs) and Demand Scrub. See also https://github.com/andikleen/mcelog
I've never seen a reason for Demand Scrub until now... If the kernel talked to the memory controller to map which pages were which banks it could cause a scrub specific high risk pages (e.g. the kernel as you say) or after X uncached accesses... Or it might be possible to segregate less trusted code to specific bank groups.
https://lackingrhoticity.blogspot.com/2015/05/how-physical-a...
Another approach might be to use memory integrity. Intel SGX has memory encryption and uses a tree of hashes to validate. It has holes against an on-machine attacker, but would defend against bit flips. https://eprint.iacr.org/2016/204.pdf
It might be hard to get the GPU to hammer memory hard enough because of the caches (http://wccftech.com/intel-skylake-gen9-graphics-architecture...). Maybe with glBufferData in DRAW mode it would force uncached access, depends if the graphics core cache snoops the address bus (coherent) or not, in which case it could be tricked. However the AMD Fusion docs certainly seem to indicate it is possible, and significantly higher read bandwidth than from the CPU (http://developer.amd.com/wordpress/media/2013/06/1004_final....). Indeed, it makes me wonder why (given the repetitive read pattern required for updating VBOs) whether the system instability people have seen before is not due to the GPU doing unintentional rowhammer.
Intel GPUs have write-back caches because they share L3 cache with CPU cores. AFAIK, other GPUs typically have write-through caches, which doesn't help against rowhammer.
"The following subscription-only content has been made available to you by an LWN subscriber."
Was that not there 15 minutes ago?
They should phrase it in a less ambiguous way.
Where is it appropriate to post a subscriber link?
Almost anywhere. Private mail, messages to project mailing lists, and blog entries are all appropriate. As long as people do not use subscriber links as a way to defeat our attempts to gain subscribers, we are happy to see them shared.