Blacksmith – Rowhammer bit flips on all DRAM devices today despite mitigations
comsec.ethz.ch
comsec.ethz.ch
Then the rate of ECC errors has to be monitored. If something is trying a rowhammer attack, it's going to cause unwanted bit flips which the ECC will correct. Normally, the ECC error rate is very low. Under attack, it should go up. So an attack should be noticeable. You might get some false alarms, but that just means it's time to replace memory.
Just the first thing that popped into my head, but: say you watch the ECC correctable error rate over time, and somehow (not so easy!) determine which process is causing those errors. You forcibly kill the process and log a message about it, and also terminate/notify processes potentially affected (say, send them a SIGBUS or something and unmap the pages containing the affected data).
I, a "clever" attacker, use this to leak out your secrets- I do my hammering juuust right so that, if some secret bit is 0, your hammering flips ECC bits, while if it's 1 your hammering doesn't affect things. Lovely little side-channel.
Universal memory encryption and authentication seems to be the only sure way out of the cycle of "attack, mitigation, attack the mitigation", and it's already starting to roll out.
Would ASLR make this harder? I assume it would because it'd be a lot harder to get to the correct memory location (?).
From physical storage to data transfer protocols, we were said you really wouldn’t bypass ECC anywhere if you wanted things to work reliably against the hard physical world. It was like a basic CS rule.
And still, there we are, still arguing that some reliability is needed into objects as critical as DRAM chips.
It's quite a lot of work, but if you have a long running embedded system, or a lot of the same, e.g. in cars, the incidence of bit flips increases.
This is it course additional to other protection mechanisms.
It won't. It's designed to counter silicon limitations to increased density, i.e. it's made to correct the errors that result from packing cells beyond the limit of error-free operation.
The extra redundancy from on-chip ECC is intended to be "consumed" by the chip itself, and since this will allow optimizing chip manufacture to denser and cheaper, it's no question at all that it will get pushed to the very limit.
There's still "classic" ECC for DDR5. 8 bits mapped to 9, terminated at the CPU which can look at things. That's what I want, need, and will buy.
P.S.: Shame on Intel for still walling off desktop CPUs from ECC. https://ark.intel.com/content/www/us/en/ark/search/featurefi...
I'd love to see actual parameters for the error correction codes, but DDR5 could pretty easily be a lot more robust than DDR4.
When you have no error correction at all, you need ridiculously high reliability. Even if these new memory cells are have a much higher error rate, if they're designed to seamlessly handle a few bits in the same row flipping then the overall reliability could skyrocket.
Edit: Oh, there's a paper from micron talking about DDR5 only having single bit correction internally. That's not as useful as it could be against attacks...
> There's still "classic" ECC for DDR5. 8 bits mapped to 9
But Single Error Correction, Dual Error Detection is not enough to prevent attacks.
Also because DDR5 uses a smaller width you actually need to map 8 bits to 10.
Well, the spec is $369, but a PDF with some early discussion (rev 0.1…) says:
DDR5 devices will implement internal Single Error Correction (SEC) ECC to improve the data integrity within the DRAM. The DRAM will use 128 data bits to compute the ECC code of 8 ECC Check Bits.
However, I have no idea whether this is what got specified in the end. 8 ECC bits on 128 data bits would be half the amount of added redundancy compared to "classical" ECC. Note also it's not SECDED, just SEC, confirming the Micron paper you mentioned.
> But Single Error Correction, Dual Error Detection is not enough to prevent attacks.
The OP suggested monitoring for a sudden increase in ECC events, which hopefully this would work for. It's not a perfect countermeasure, just a statistical one, but I'll take what I can get...
> Also because DDR5 uses a smaller width you actually need to map 8 bits to 10.
Now that kinda makes it better, doesn't it? :)
Which is a shame because it would only cost 1 more bit to do SECDED.
> Now that kinda makes it better, doesn't it? :)
Very slightly, but 32->40 isn't much better than 64->72.
SECDED requires 39. I don't know if the 40th bit is used for anything? It would certainly be possible to use the 16 extra bits per burst to add a second layer of SECDED. Or an extra layer of triple error detection, if I'm remembering right. That would be really good against rowhammer.
FWIW this is completely up to the CPU. "Classical" ECC is completely transparent to the DIMM, it just remembers and returns what the CPU gave to it.
32 → 40 has already been around for quite some time (particularly when the DDR controller only has 64 data lines but you still want ECC you can sometimes do 32+8 at half performance & max capacity) but I don't see anything beyond SECDED in the datasheets... (looked through some Freescale/NXP PowerPC bits)
The problem is that this assumes that a probability of flipping individual bit is independent from others. But this may not be the case. And if so rowhummer is possible.
It’s not end to end ECC as in it doesn’t prevent flips that happen on the bus or in CPU cache but it does prevent single bit errors on DRAM.
Bit flips that are getting more and more common as the RAM cells are getting tinier and tinier, the stored charges ever smaller and smaller, and thus susceptible to flipping…
I agree it's meant to counter limitations due to increased density, but it should catch this to an extent also as this error is induced on-package right, not during transmission. Or am I mistaken?
> What if I have ECC-capable DIMMs?
> Previous work showed that due to the large number of bit flips in current DDR4 devices, ECC cannot provide complete protection against Rowhammer but makes exploitation harder.
Does that mean we need an extended ECC to deal with critical systems that require additional robustness?
Until someone starts building chips with fake ECC to boost profit.
> What if I have ECC-capable DIMMs? > Previous work[1] showed that due to the large number of bit flips in current DDR4 devices, ECC cannot provide complete protection against Rowhammer but makes exploitation harder.
1. https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=8835222
The fact that user space code can cause bit flips in your RAM is a hardware bug. I'd love to see this code in memory testers like memtest86 so I could send the RAM back if it ever caught a problem like this.
I guess it shows just how close to the edge of not working our modern computing environment is.
And I have no idea if the internal "ECC" of standard DDR5 helps or not. It is not intended for regular ECC level of reliability anyway. (And I have seen discussion about likely bitflips detected in crash dumps of M1 Pro devices)
So, as much as I would like to return defective devices, I would probably be left with no computer, no smartphone, etc.
> “DDR4 systems with ECC will likely be more exploitable, after reverse-engineering the ECC functions,” researchers Razavi and Jattke said
> What if I have ECC-capable DIMMs? Previous work [1] showed that due to the large number of bit flips in current DDR4 devices, ECC cannot provide complete protection against Rowhammer but makes exploitation harder.
[0] https://cs.vu.nl/~lcr220/ecc/ecc-rh-paper-eccploit-press-pre...
[1] https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=8835222
Yes, and this is actually the problem. Without safe hardware, it is almost impossible to write safe software.
(Have done this a few times. Was hesitant initially because of the low price and used nature, but when the first stick passed MemTest86+ including RH perfectly, I was convinced. Half a dozen more sticks later and still good. But maybe they've started to run out of the good old stuff now...)
That's demonstrably false. I remember rowhammer being thought of and demonstrated on DDR3 modules.
The tl;dr is that everything 2009 and older is good, 2010-2011 is possibly bad, 2012 and newer is almost certainly bad.
Software that attempts a hostile rowhammer attack aims to be unsafe by definition. It's not a matter of trying and failing to be safe. Rather, we're seeing that it's very difficult to build hardware and software that prevents hostile code from running amuck: if the attacker can run hostile code on your computer at all, you have already lost. That's bad news for cloud hosts and confirms what we already knew about web scripting, but doesn't seem all THAT bad for the rest of us.
Of course it's possible to build computers with no DRAM at all (i.e. only SRAM, which costs more). That is not economically sane for big servers, but might be the right thing to do for security sensitive components.
If that is true, vs just the usual "you don't need to know" trash common these days, its crazy. A pretty good picture can be built up with just software timing analysis, but its hardly rocket science to hook a logic analyzer/scope up and determine bit swizzling and page/bank interleave. Plus, many of the more open BIOS vendors have tended to provide page/row/controller/socket interleave options for quite a while, partially because it can mean a 5-10% uplift for some applications to have it set one way or the other. Its been one of those how do you tune your memory system options for a couple decades now.
IDK though. Maybe there's a way. I've worked on code that uses padding around a data structure to ensure it has it's own cache line. Maybe if you allocate a large enough contiguous block you'll be OK?
In other words, the whole industry sits on a bed of lies at this point and it's only because government is technically incompetent that we haven't seen the world's biggest class-action.
That's a pretty bold claim - if it's accurate, can you point at the government regulation that precludes a class action lawsuit? I would assume it's something in the T&C or EULA for the hardware rather than a regulation?
If users knew better, we'd have the equivalent of another Pentium FDIV bug.
All computed results from data science must include steps and code to verify locally.
It calls into question privacy on federated networks and crypto networks; any node can be manipulated locally to change payload outputs on delivery, reveal secrets, disrupt workloads.
This makes sense to a lot of folks in computer engineering and physics, versus abstract software. No physical theory I know of offers any guarantee our arbitrary computing machines will ever be securable. We put fart pipes on Hondas.
Science proves it’s titillating smoke and mirrors once again. Still waiting for nuclear rocket cars.
I think this proves further as well why general computing chips need to be replaced with workload specific designs, where the anticipated inputs are well known and no vague logic paths to intentionally allow software monkey patching ever ship.
Perhaps we fail at pricing security into the value of a company, or maybe that’s what risk appetite is about.
This is worsened by the fact that it's very hard for laypeople to assess the security of a specific application and that, by now, "cyberattack" has become common enough that it's easily accepted as an excuse.
The market just yawns at this stuff, until it gets fragged. Then it forgets and the cycle repeats.
Not sure about that. All the security standards want me to run software written in an unsafe language as root on every device, intentionally parsing malicious inputs continuously.
That’s not making anything safer.
You can run in dedicated tenancy where you have the whole machine or a metal configuration where you also have the whole machine.
(And those who employed hardware mitigations wouldn't have a problem with user:hardware ratios, would they?)
Is this an individual stick bug, or a design bug akin to Spectre? Given that they tested 40 different sticks and seemed to find effective patterns for all of them make me think it's the latter. In which case, just return all your computers?
The article states that ECC isn't sufficient to solve Rowhammer, but it does make the attack harder. If you are looking for a hardware solution it is your only option currently, even if it is imperfect.
Google -> Alphabet
Facebook -> Meta
Meta, Apple, Amazon, Alphabet, Netflix. MAAAN.
Every now and again I worry I've grown up. Then I realize there is no grown up, just a Matroyshka doll of more and more eclectic subtexts
M - Meta A - Apple N - Netflix G - Google A - Amazon
Apple Microsoft Amazon Netflix Alphabet Meta
This is not a tenable solution.
You seem to know more about this than I do. Where might one go to read more about this? Are there whitepapers, design docs, other resources which could be used to assert positively that the industry used to consider this a defect?
I believe newer versions of MemTest86+ have Rowhammer tests but it's disabled by default.
See my previous comment on this matter:
Most of the higher end hynix/micron/samsung sticks I've tried do not fail at JEDEC or XMP after 7+ passes.
Except that the distro's are still shipping an older version due to license issues (or something like that AFAIK).
How do you run a memory tester after you have returned all CPUs after Meltdown and all the subsequent microarchicture bugs? I'd call all modern CPUs faulty products. Some more than others.
Computing moves in cycles:
- 2000s: gigahertz race on each core
- 2010s: increase core counts, multicore everything
- 2020s: back to core-efficiency and increasing per-core performance. M1 is already leading this charge, but is obviously a mismatch for a DC environment.
AMD and Intel need to adjust, or face extinction. Its not just about pushing ultra-high per-core performance (they're both good at this); its about pushing for more efficiency, so per-blade density in a DC can be pushed higher in the face of more, smaller individual computers. If they don't evolve, AWS will for them [1].
Interesting that AWS has been mitigating for side channel attacks since before they became a big news item. Curious about Azure and GCP's stance on this
Or, maybe, they just thought the lack of deterministic performance created billing/accounting/customer service problems. (One hyperthread can just about completely starve in many circumstances).
"Are We Susceptible to Rowhammer? An End-to-End Methodology for Cloud Providers"
https://arxiv.org/pdf/2003.04498.pdf
Although their answer in this paper was diplomatic, my interpretation is that they confirm it as a problem. Their conclusion was it would not be as bad it was considered at the time. To be revisited on the context of this more recent work.
Edit: Adding main reference
"BLACKSMITH: Scalable Rowhammering in the Frequency Domain" https://comsec.ethz.ch/wp-content/files/blacksmith_sp22.pdf
You pick a VPS because you want to save costs and share cores.
So it’s not so much AWS choosing as you the customer choosing.
I ran the numbers back in 2015 and persuaded the previous employer to go for dedicated tenancy with all performance-critical and privacy-sensitive workloads. I was effectively hedging against unknown but practically guaranteed cross-VM attacks to pacify a paranoid regulator.
Then Rowhammer happened. Less than a day later, our contact with the regulator comes to me asking how it affects us. Being able to answer - with absolute confidence - that it did not, was one of the proudest moments of my career. And the turnaround from "regulator comes asking awkward questions" to "regulator is happy and sees no reason to ask again" of less than 48 hours must be some kind of record too.
The looming thread of side-channel attacks on SMT systems has been known since well, before we had SMT systems (because it also can apply to Co-Processor, and non SMT multi core systems).
The difference between back then and today is "we believe it's possible but haven't found a way yet" and "there are multiple known ways", as well as it being wide spread known instead of just in some communities.
The reason we still shipped the problematic CPU's is because improvement in perf. and as such competitiveness and revenue on the short term where more important.
There also was a shift in what people expect from security and which attack vectors are relevant. For example user applications a user installed where generally trusted as much as the user. While today we increasingly move to not trusting any applications even if they are installed by a trusted user and produced by a trusted third party. Similar running arbitrary untrusted native code from multiple untrusted users and "upholding" side-channel protection wasn't often an important requirement in the past.
I think the big difference is that we have attacks with high bandwidths of exfiltrated information-- instead of very slow, difficult to control leaks.
Sure, but this is worse than that. This is "your online poker game client can gets access to your web browser's bank account session info."
We need process isolation within a single machine, or else we are kinda screwed as a field.
But this is an electrical problem. Interference is a huge issue with any engineering that involves physical things and these kinds of attacks are just interference problems. This issue is no different from a microwave knocking out your Wi-Fi. These attacks have become possible because the acceptable interference threshold that chip makers have been using has turned out to be too low.
How do you fix interference problems? First, you choose a new threshold of acceptable interference and then you engineer better isolation, you lower density, and/or you switch technology.
We could make shared computing complete safe tomorrow if we wanted to so I think calling this the end to shared computing is quite alarmist. The issue is that we collectively want to both have the cake and eat it too: we currently have a certain cost-to-compute ratio that we have become accustomed to and we don't want to compromise that. We're basically buying time until we can invent a new technology that can achieve the same density without the same level of interference.
So maybe "screwed as a field" is not an exaggeration if the field is butt computing.
Think about someone trying to do basic ETL. Like having a tabular file and summing it or something. Don’t use Excel, we say, stand up a $4 million Spark and AWS architecture with seven hundred pitfalls that can let bored Russians take over your whole network as if they were going to the dry cleaners because remember, you just might be Google someday. That’s where we are. It’s been a complete industry failure for a decade and it’s only getting worse. Accelerating, even: now you need some operationally-terrifying Kubernetes to even be at the table, and then as an industry we (rightly) say running this stuff ourselves is too hard, so pay Amazon to do it rather than even ponder if we have settled on the right approach.
Tada: Humanity just lost computing to three companies. We very likely aren’t getting it back.
There are probably 5,000 people doing this work who can adequately secure such a system and make it mostly impermeable. Where middle computing is royally screwed is that nearly all of them work in San Francisco or its clones abroad. So then you get “best practice” blog posts and industry think pieces and the lowest bidder ties them together into something resembling a competent computing system. That’s been the state of the art since 2004 everywhere except Santa Clara county.
With the exception of some areas in the IC and DoD, I just described the entirety of US government IT. That ETL example? It’s actually real and underpins a small part of Medicare across several government contracts. Because the tools the valley exports are all they’ve got, and we sure love building systems with massive footguns, and then shaming organizations publicly for missing item #543 on the tribal “secure your computing system” checklist and shooting themselves.
The entire industry must change, top to bottom, but just like climate change, again, that’s a nonstarter. Posix and the Web are not the path forward and I hope I live to see the industry figure that out. I’m increasingly skeptical. The good news is my hometown might flood into the sea first, sparing me from considering in my last moments that every argument I’ve ever made in this profession has fallen on deaf ears and that everyone has to derive our industry’s peril from first principles for themselves.
I think it will come down to what you are willing to call 'individual' processors. But actually having physically distinct memory seems like a lot of overhead for attacks that won't matter for 90% of users. Also I would think that the on-board ECC of DDR5 would protect it against these types of attacks.
The trend is yet to come. My statement is that, if AMD/Intel doesn't adapt, Amazon has the hardware investment to leave them behind, just like Apple did.
But to be clear on two points: They will probably adapt. And Amazon/etc will probably never leave them behind fully. DCs, especially public cloud, are not all-or-nothing like Apple's Mac Lineup is.
Then the question follows, why would they want something Intel/AMD isn't offering right now? The trend is System-on-Chip. Beyond Security (this isn't the last electrical interference/speculative execution-like attack we'll see). SoCs are easier to service (easier != cheaper. holistic replacement versus per-component debugging. servers are cattle, not pets). Denser. More vertically integrated. Capable of far higher IO performance. Lots of benefits.
Mega-servers with 256 cores and 4 terabytes of memory still have a huge place in all DCs; but not when multiple untrusted workloads are running simultaneously. They're not for EC2/Fargate/Lambda/etc; they're for S3. Highly managed, trusted workloads.
And indeed to my understanding spectre, meltdown, rowhammer and similar attacks are not an issue there.
https://en.wikipedia.org/wiki/Z/Architecture
I wonder when more features from the mainframe cross pollinate Intel amd arm CPU architectures.
A lot of the talk of mainframe levels of security is specious at best.
https://www.suse.com/de-de/support/kb/doc/?id=000019105
Meltdown didn't affect z nor amd.
https://www.zdnet.com/article/meltdown-spectre-ibm-preps-fir...
https://wiki.debian.org/DebianSecurity/SpectreMeltdown#Syste...
https://tenfourfox.blogspot.com/2018/01/actual-field-testing...
If I understand correctly, Homomorphic Encryption aims to solve for these kinds of attacks (although presumably the computations are more expensive and programs must be restructured to use HE primitives?). https://en.wikipedia.org/wiki/Homomorphic_encryption
EDIT: Why the downvotes? Am I mistaken?
Not to mention, to avoid side-channel attacks, an HE scheme still needs to always run the longest possible sequence of operations regardless of input (otherwise, information about the input data is leaked through timing). So an HE version of a quick-sort scheme would always run in O(n^2), otherwise it would leak details about the contents of the list. In some cases it would even have to run in the same amount of time regardless of the size of the list, to avoid leaking information about that.
This is still phenomenally slow, of course.
It works a little bit like this: you write your algorithm using only HE computing primitives, then you encrypt it and the input data and send those to the HE runtime, who you don't trust, and who DOESN'T have any way to decrypt your message. The HE runtime executes your encrypted program and sends back the still encrypted output. Once you receive your output, you decrypt it using the key you have kept secret.
I'm not sure what you mean. This is how HE works, at a high level. The whole point is that you're doing public operations on secret data on a machine that can never know what the data was, nor what response it gave you. Furthermore, the point of HE is also that you know that your data remained secret regardless of how the HE runtime was implemented, as you just send it encrypted information.
Making it more efficient by giving it access to the decrypted data would defeat the whole point. Making it more efficient by optimizing based on the actual data would allow your data to leak through timing side-channels, again defeating the whole point.
Basically, HE refers to computing schemes where DECRYPT(ALGORITHM(ENCRYPT(X)) == ALGORITHM(X). You send [ALGORITHM, ENCRYPT(X)] to the runtime, and it gives you Y. DECRYPT(Y) is then the result of your computation.
As far as we know, this is indeed hopelessly slow, and there is no guarantee that we will ever find fast secure HE primitives.
And yeah, this is just as inefficient as you'd think.
Firstly the computations on encrypted data just take a LOT longer, especially for Fully Homomorphic encryption. With a single 64 bit addition taking microseconds.
Secondly, the FHE code cannot branch based on data. If it could, it would know something about the data, and it wouldn't be proper encryption. This means an if statement becomes "calculate both branches, and throw away the result you don't need". Similarly, a for loop becomes "give me an upper bound of how long this will run", then "loop for exactly the upper bound number of times, if the loop is 'done' early, just throw away the result of the remaining operations".
FHE is really cool, and has its uses in situations where you want to cooperate without needing to trust, but it is stupidly inefficient. (Things get really cool if two parties want to cooperate and the computing party can e.g. branch based on their local unencrypted data)
But it's likely things like different core counts, frequency performance curves, perhaps more or different memory/IO controller counts, perhaps instruction extensions of some particular interest (I'd wager BF16 was available in data center SKUs long before consumer SKUs), etc.
It's nothing radical, but if you're going to buy an awful lot of them and want something slightly custom, Intel will definitely do that for you.
Cloudflare mentioned they use a custom Intel SKU on their blog, and they're not as big as AWS.
Your take sounds accurate for non-FAANGs (and a couple of those). For Google, not so much. There is a lot of custom acceleration and such that isn’t available to any other customer, and a nearby commenter is right, every detail about them is locked behind a very strong NDA. I’ve heard thirdhand AWS does indeed have custom instructions in the virtualization extensions stuff, but don’t know if that’s true.
Right you are, Ken.
Depends on how many you order and how reasonable the customization is. Like, what if I wanted a 10,000 38 core Xeons with a 2.5GHz Base clock and 3.5GHz Turbo clock? Such a part doesn't exist, yet sits smack dab between two existing Models, the Xeon Platnium 8368 and 8368Q. Assuming yields are already good enough, that might be the sort of thing a CPU maker will do. (Might need more than 10k units though, IDK.)
Intel does not want to be competing with an Amazon Basics processor.
"The custom Epyc 7003 part that Facebook has commissioned from AMD has 36 cores out of its potential 64 cores activated, and with a bunch of optimizations for Facebook applications at the web, database, and inference tiers of its application stack."
While this seems to be mainly about Power/Heat/Perf. optimizing for typical "cloud" workloads it might also have design decisions related to this problem. We will see, but they look interesting anyway.
I'm curious about the eventual end-game of security in this space. Take the 64 individual processors in your example, give each one their own independent memory bus to their own ram chip, isolate them from each other as much as possible. What else can be done, if a malicious process on processor Z has to go all the way to disk to try to get back at data working on processor J, is that as maximally secure as it can be without being in a completely separate chassis with only network access to the other device?
Essentially, a board with an array of "mini computers", 1-4 cores bussed to their own physically and logically isolated memory pool. Storage is oftentimes already networked in public cloud, and seemingly less susceptible to these kinds of attacks, so probably little change is necessary there.
Its also likely we'll see vendors produce more "card computers", like the raspberry pi, but oriented at a server environment. At massive scale like what AWS runs at, these are less compelling, but they do have advantages; the biggest being, maintenance is a cinch, and the blast radius on one failing is isolated to just the workload on that card. I imagine these will be more interesting for current VPS providers like DigitalOcean, who also want to offer dedicated hosts with smaller core/memory counts. Even RPis today are super compelling; a $50 multicore computer that's easy to manufacture and replace, with density high enough to fit 100+ inside a 4U? Just need a bit more core perf and ARM adoption (though its easy to see Intel also hoping on this train with their long-term investment in pro/consumer NUCs).
Writing is on the wall. The tipping point will be an industry thought leader saying, in effect, "you should be using dedicated hosts, because at this point physics is our enemy and we're not winning, by the way here's a new dedicated host instance type that's perfectly suited for horizontally scaled workloads, its 20% more expensive than virtual hosts of the same size, byeeee".
I have server application capable of utilizing many threads and thousands requests/s. You think I will deploy it on tiny 2 core CPU? No thank you, it currently runs on powerful dedicated server from Hetzner where I control everything.
>"AMD and Intel need to adjust, or face extinction ....
If they don't evolve, AWS will for them..."
Sounds like pontification / FUD.
Sounds like you didn't fully grasp what the grandparent was writing about.
>> Eventually, it will be obvious that running shared workloads on a single piece of physical hardware has fundamentally unremediatable security implications
> I have server application capable of utilizing many threads and thousands requests/s. You think I will deploy it on tiny 2 core CPU? No thank you, it currently runs on powerful dedicated server from Hetzner where I control everything.
Your use case isn't covered here, you're neither a cloud nor VPS customer.
There will always be individuals and companies looking for dedicated hosting solutions, or just hosting it themselves.
However, for the rest of the world that wants to run their stuff in the cloud that's where Amazon's and other companies investments in ARM come into play. That's where we'll see disruption, the extinction the grandparent was talking about.
Unfortunately it will likely have negative performance implications for multi-tenant work loads.
That being said, and even if I consider that a quite different subject, I agree that the current efficiency story of Intel is not very good, but I hope they will improve in a not so far future. The dev lifecycle of CPUs is quite long and it seems to be an obvious target. I suspect they will be forced to improve their efficiency, because that's actually were the performance potential is today (the current dissipation level of their last desktop CPUs is not reasonable, and prevents scalability). And trying to lower the core count also can yield to high consumption, e.g. if you want your performance back by increasing the frequency. Wide and "slow" is needed, and it is harder to increase the internal number of execution units per core and have them actually used, than to increase the core count -- plus ironically one way to do that is through HT, which goes against your wish to share less hardware. (Now if you compare their P-core and E-core in Alder Lake the story is more complicated, but their marketing figures seem very strange, so I won't conclude anything for now. The current instances of P-core we have is for that weird desktop market with unreasonable high TDP anyway.)
Now if you really want miniaturized individual computers that would not be shared at all, I'm not sure the market will actually go into that direction because big systems will continue to be needed (and clusters are a niche mostly for HPC), and I'm not sure a "slightly more secure on smaller systems" market would be interesting enough. Esp. in an era of chip shortage. And also because it would still be bigger than a shared equivalent system make with the same techs... But if that's really a niche that has to be addressed, I suspect it would mostly be a matter for Intel to create new small and slower SKUs ("slower" compared to their desktops insanities) -- they even kind of have that already, but yes the physical miniaturization aspect is not handled yet -- that does not really depends that much on the cores, though. And even in those computers, I'm not sure there would be much demand for very low core counts. The threading of pretty much all workloads tends to increase, nowadays.
One last point, after the e.g. Pentium 4 fiasco nobody really left the IPC race. AMD had some difficulties when trying "weird" ideas (part of which were because of their marketing communication), and then again a completely new design from scratch to market takes time. In general there was a pause of performance growth around 2016 for a few years, and that was mostly Intel having process problems and the rest of the industry catching up (and then overtaking them).
For me cloud computing is just where the best pay is. I do not at all see it as the future of computing.
One reason is ML will help us realize we write code we don’t need; so much of it is syntax sugar for business specific needs; infra, security… it’ll be realized cloud software is solving unemployment not technical problems of value. That many issues with software back in the day were lower quality networks, and consumer hardware. I mean any phone can abstract metadata from any one users amount of behavior, we do it in a DC because that’s where the jobs are. Chip manufacturing will include ML normalized logic for specific application.
LAN IOT will improve and we’ll realize the Metaverse can be implemented with a local client and AI generated art, on a mobile GPUs power in a few years. Middle men like Zuckerberg face the most uncertain future. He failed to diversify as well as Bezos, Newell, and others.
IMO, Valve is a serious threat with Steamdeck; an open IOT brain in a kid friendly form factor could be the new cigarette. Even Apple may have to take them seriously. My kids iPads need replacing soon; a flat glass slab with no interactive controls, requiring another $800+ machine to develop on, bloated development tools, fees, and a bunch of cloud logins, are not going to motivate kids to feel creative.
If you have a process running on your machine can it use this to get root? Read Keys?
It looks like they ran their process for 12 hours to do the flipping.
And if your flipping your process's memory for that long, what are the chances you are next to sensitive memory for another process? It seems bad, but it seems like if your randomly flipping bits in memory the system will likely crash.
Yes. There's a BlackHat talk on that[1]. DoS is a big issue on shared hardware too. You can crash other processes, the kernel and the hypervisor.
> And if your flipping your process's memory for that long, what are the chances you are next to sensitive memory for another process?
You don't need to leave it up to chance. There are apparently ways you can control it and of course you can just spray everything.
1. https://www.blackhat.com/docs/us-15/materials/us-15-Seaborn-...
2. https://github.com/google/rowhammer-test/blob/master/rowhamm...
My (probably slightly naive) understanding is the OS is allocating memory, acting like a memory cop so processes don't overlap and swapping when needed and so forth. It seems like these hardware errors might be mitigated by the OS.
I'm in a bit over my head in the pdf, but it explains: "Our kernel privilege escalation works by using row hammering to induce a bit flip in a page table entry (PTE) that causes the PTE to point to a physical page containing a page table of the attacking process. This gives the attacking process readwrite access to one of its own page tables, and hence to all of physical memory."
I'm still a little fuzzy on how they get the right location in the page table without hosing the system, but this gives me enough of a gist. Thanks!
You don't need to have access to the memory you want to attack, only the ability to cause something to read from memory that physically neighbors the memory you want to attack.
Think of page tables like a big array in kernel memory (I'm being imprecise.) The entries in this table are mappings from your virtual addresses to physical addresses. You can cause the kernel to read your at least part (or all of?) your PTEs, which means you can cause it to flip bits in neighboring physical cells, which means you can modify the PTEs. TL;DR: If you can cause the kernel to read some specific memory then you can modify memory near it. In this case, you can modify the page table entries, allocating memory that isn't yours to you (or doing other funny things.)
A simpler to explain but probably not realistic explanation: Imagine if your UID is stored right next to your GID in some kernel structure, if you could cause the kernel to repeatedly read your GID then you could cause bit flips in your UID, which would change which user your user...
Is there something I'm missing here? What are the realistic attacks that this vulnerability allows?
Which keys? OpenBSD / OpenSSH for example does now (since 2019?) encrypt SSH keys in RAM to prevent rowhammer like attacks. They're mixing the key with a huge "pre key" and an attacker would need to side-channel attack the entire pre key to be able to read the real SSH keys.
So it's not as if there was nothing that could be done to guard against this.
It may require changing lots of software/libraries but at least the security-conscious ones already started, without waiting for this "blacksmith" attack.
Combine that with a lack of knowledge of physical memory addresses and inability to have much control over memory layouts, and it really gets tricky to gain privileges outside a lab environment in a reasonable short time.
Remember that flipping bits at random will almost certainly kernel panic the machine before it gives you root access.
I'm sure a determined attacker could do it though.
Not quite. It doesn't require the ability to write to the neighboring cells, just read from them.
On every boot, there is a brand new kernel and C library:
reordering libraries: done
reorder_kernel: kernel relinking done
ASLR has been compromised in the past, so this likely isn't completely secure.At Alpha Microsystems, about 1982, I was in charge of the diagnostic programming group. At that time failing memory boards were expensive and customers would not be happy with such components.
I wrote a memory testing diagnostic that was based on knowing exactly how addresses were mapped to cell locations so I could try to provoke such failures.
Chip manufacturers were aware of this problem which is why they scrambled the addresses.
Potential vendors, Motorola et al, were required to provide mapping information before we would consider their chips.
Now I'm curious to know what such mapping looks like with modern memory chips.
1976-1985 My first job was at Basic Four Corporation. I got in as a test technician, assembling small refrigerator-sized mini-computers, giving them their first tests and swapping components until they passed. I soon learned how to use the machine-language assembler and started writing small programs to help me determine which components were failing without swapping and hoping. Within a few months I was testing more than 2x the number of machines the other techs were testing. Management noticed and soon the other techs were getting up to speed.
At this point management pretty-much turned me loose. I moved up to diagnosing and repairing failed components (8"x11" pcbs). My understanding of programming and digital circuits allowed me to write small programs that could be used to "light up" specific circuits on the board making it easy to poke around with a scope to see where things went wrong. Again this was a huge productivity boost and the technique was propagated to other techs. A couple years later I went back for a while part-time as a consultant and wrote my first DSL for techs to use.
Next I talked my way into the firmware development group and worked on firmware for tape, disk and other devices. This is the period of time when microprocessors were being incorporated into everything and my experience with the Micro-68 put me in a good position to participate. I also got to write microcode for a 2901-based cpu that was in development.
And then, somehow, perhaps at a user's group, I learned about Alpha Microsystems.
When I first visited Alpha Microsystems their idea of "burning in" pcbs consisted of putting them in a powered backplane, in a wooden box, with a lightbulb, where they sat for some period of time.
Basic Four had serious testing which included putting entire computers into temperature-cycling ovens where they ran tests for 24 hours. That knocked out a pretty high percentage of boards. After I told them what they were missing Alpha Microsystems hired me to improve their process.
For the next several years I participated in creating the flow of production and testing. A department evolved to handle the hardware side of things and I became head of a diagnostic programming department which grew to, variously, between 6 and 10 programmers. After that department was functioning and had someone who could step up I transferred into the operating systems group. I was one of only three people allowed to work on the operating system code, let alone even see it since it was held as a trade secret.
During my last year at Alpha Microsystems a brilliant programmer I had hired introduced me to the then just-released Structure And Interpretation Of Computer Programming, a new textbook for students at MIT. That book's use of the language Scheme introduced me to first-class functions, closures, and many other concepts which found their way into popular programming languages decades later. The SICP had all the information one needed to create a Scheme interpreter. I wrote one using 68000 assembler so I could run the sample code.
i love programming so much i've never quit
taught fullstack at ucla BC (before covid)
did a large frontend project with vue the past year
i'll die with a keyboard under my fingers!
Yowch. Like PC-DOS 1.0 all over again.
You need to bear in mind how limited disk storage was in the 70s and early 80s. When I got to AM systems were booting from 8" floppy disks. That gave you about 500kb, iirc.
The file system had "directories" in this form:
[nnn,nnn]
Where, again iirc, nnn was 8 bits, so 11 111 111, 377 for example. Only 16 bits were required to specify the location of a file.
Filenames, btw, were six characters.
Sorry for that offtopic rambling, I am little bit enthusiastic every time, if I see something TOPS10 related.
the western digital chipset that powered the first am systems was a five-chip 40-pin set connected by a private cmos bus ... probing with a scope was amusing
the instruction set was basically identical to the pdp-11
It also brings to mind that Simpsons scene: "Stop stop! He's already dead!"
1. https://www.youtube.com/watch?v=AJBmIaUneB0&list=PL5Q2soXY2Z...
(As for the possibility of such, there is already an x86 to x86 JIT compiler that increases performance)
Any speculation around memory accesses will yield this.
Also, In dougallj's code [1] the zero of registers should be superfluous so it is assumed the function below is needed to make the experiments run stably by claiming ownership of the registers as part of a general anti-speculation security mechanism.
static int add_prep(uint32_t *ibuf, int instr_type);
The M1 explainer [2] has lot of interesting ideas like this contained inside it.
[1] https://gist.github.com/dougallj/5bafb113492047c865c0c8cfbc9...
Which is only possible because the DRAM has limited memory for recently-accessed rows.
When is a company going to put out chips that have the access count stored inside the row? It's the most obvious way to do it and makes this entire class of attack impossible.
Edit: Okay, reading the paper more apparently LPDDR5 has something similar to this. Why is LPDDR so divergent from normal DDR?
If it comes you'll hear about "apps" that freeze every other browser process while you do online banking.
That said, proactively blocking JS is great for many reasons, and you can whitelist sites you sort of trust.
This is a pretty laissez-faire attitude with regards to security-critical exploits. You might be the one of the first ones affected. After all, someone has to be first one before they can post it on HN.
In a more general sense, it's a bad idea to rely on someone else as your primary alert. Whether it is a HN post, news article, whatever.
(I know that you aren't quite saying to use it as a primary alert, but a cursory read of your comment implies it, so it's worth discussing.)
If you are regularly the fist target of zero-days your threat model should be stricter yes. The other 7.9 billion people in the world can afford to be a little less strict.
It did come to that, years ago: https://github.com/IAIK/rowhammerjs
I mean, the last fappening was done by phishing, as I recall. Are you saying another one, done by rowhammer, wouldn't be a problem?
I suspect most folks have stuff they'd rather not make public. Credit card numbers, at the very least. This affects us all.
Similarly the day-to-day ransomwares probably cost more than securing browsers will.
Something about CPU firmware and its inability to fix some aspect of memory bank controllers.
I've been browsing with Javascript off-by-default for years and recommend it regardless.
Sure, perhaps a site I trusted and whitelisted will try to hack me, but I feel it's far less likely than some random typo-squatting site or advertising-broker logic.
Not at all. It is another one of the Chicken Little's security theater that sound sexy and sophisticated but in practice mean nothing.
If you are concerned about this issue you should first ensure that you never type "npm install" or copy something off Stack Overflow into your code.
https://github.com/IAIK/rowhammerjs
This has existed for years and its use has been observed in the wild. RowHammer isn't new.
You pulling 15Gb of Nodejs modules in your app that you have no idea about because you think someone, somewhere ensured that they are "correct" importing a supply chain attacks works every single time.
> “DDR4 systems with ECC will likely be more exploitable, after reverse-engineering the ECC functions,” researchers Razavi and Jattke said.
> ECC cannot provide complete protection against Rowhammer but makes exploitation harder.
edit: I looked at some linked papers and I see similar quotes, though not that one.
edit2:
https://arstechnica.com/gadgets/2021/11/ddr4-memory-is-even-...
> “DDR4 systems with ECC will likely be more exploitable, after reverse-engineering the ECC functions,” researchers Razavi and Jattke said.
OK, so not more exploitable relative to not-ecc RAM, just relative to ecc ram pre-RE.
The ECC code word is bigger so it is a larger target, but you have to flip multiple bits to cause pain. If you have 2 bit detect, you need to flip three bits to get something that corrects to a different value.
Local DoS? Can you elaborate on this?
"...Are there any DIMMs that are safe?
We did not find any DIMMs that are completely safe. According to our data, some DIMMs are more vulnerable to our new Rowhammer patterns than others.
Which implications do these new results have for me?
Triggering bit flips has become more easy on current DDR4 devices, which facilitates attacks. As DRAM devices in the wild cannot be updated, they will remain vulnerable for many years.
How can I check whether my DRAM is vulnerable?
The code of our Blacksmith Rowhammer fuzzer, which you can use to assess your DRAM device for bit flips, is available on GitHub. We also have an early FPGA version of Blacksmith, and we are working with Google to fully integrate it into an open-source FPGA Rowhammer-testing platform.
Why hasn’t JEDEC fixed this issue yet?
A very good question! By now we know, thanks to a better understanding, that solving Rowhammer is hard but not impossible. We believe that there is a lot of bureaucracy involved inside JEDEC that makes it very difficult.
What if I have ECC-capable DIMMs?
Previous work showed that due to the large number of bit flips in current DDR4 devices, ECC cannot provide complete protection against Rowhammer but makes exploitation harder. What if my system runs with a double refresh rate? Besides an increased performance overhead and power consumption, previous work (e.g., Mutlu et al. and Frigo et al.) showed that doubling the refresh rate is a weak solution not providing complete protection.
Why did you anonymize the name of the memory vendors?
We were forced to anonymize the DRAM vendors of our evaluation. If you are a researcher, please get in touch with us to receive more information. ..."
By coincidence, I’ve been bluescreening with ram related error codes for the last 2 days haha
"Drive-by Rowhammer attack uses GPU to compromise an Android phone" [2018]
https://news.ycombinator.com/item?id=16984663
"Inducing Rowhammer Faults through Network Requests"
$value XOR f(memory address, random number etched onto chip)Even then, memory vendors tend to want to compete on frequency and access timings, which means doing any additional work not strictly required by the JEDEC standard will make their product appear worse than competing products, so I doubt they will want to do that.
Plus a similar technique could actually be done by the CPU's memory controller to similar effect, and historically DRAM design has favored pushing things to the controller when possible.
1. https://www.blackhat.com/docs/us-15/materials/us-15-Seaborn-...
I highly recommend that series by the way.[1] It's not just an architecture class because it's actually pretty up to date and highly focused on security.
1. https://www.youtube.com/watch?v=AJBmIaUneB0&list=PL5Q2soXY2Z...
TRRespass: Exploiting the Many Sides of Target Row Refresh
https://arxiv.org/pdf/2004.01807.pdf
Memory controller mitigation of RowHammer can work pretty well, if one actually has it turned on. Which is unfortunately rarely the case even in 2021.
It’s cheaper to buy a dual Xeon 5660 (six CPU times two chipset times two for SMT) and ply it with 244GB of DDR3-1600 to get an overclocked 4.4Ghz than it is to buy anything DDR4.
This extends well beyond software.
Next time you smell "gas", give a thought to the 300 souls of the New London School.
https://en.wikipedia.org/wiki/New_London_School_explosion
Safety standards are written in blood.
>Until there are victims, threats tend to be ignored.
If you can't convince someone to address a threat, you need to get better at convincing, rather than creating victims to prove a point.
History strongly suggests it is not.
In large part because non-manifested risks, or not fully manifested risks tend to be dismissed or belittled, most especially when that denial is presently remunerative. Vaccination and non-pharmaceutical interventions against global pandemics, an economy based on massive surveillance by states / enterprises / other actors, global warming, CFCs and ozone depletion, use of lead (in fuels, paint, glazes, toys, and other materials), asbestos, tobacco, car and highway safety, racism and ethnic favouritism, industrial safety, mercury use, plastics, unoderised natural gas, mine safety, electrical codes, fire codes in buildings, ...
In the infosec / cybersecurity world we've at least the advantage of being able to provide proofs of concept which move a given risk from the theoretical to the manifest. Even then, effective reduction or mitigation takes far too long.
If there's a correction, I'd suggest that the onus be placed not on the advocates but the deniers. A marked penalty for advocating a position which turns out to be untrue with credible basis for knowing that it is untrue is one possibility that's occurred to me. How that would play out I'm unsure, but I'd be willing to explore the question.
You're advocating ought.
I'm describing is.
I'm proposing ethical ways, wishing harm on others is unethical. You're free to make your own choice.
I've already stated my case as clearly as I can, and why my preferred choice would be. Your response indicates that's insufficient to convey my intended meaning to you. Which is an interesting case of communications failure itself, come to think of it.
Again, PoC is the widely-accepted alternative to general catastrophe. Even that is very often insufficient. In a "less harm* calculus, an exploit in the wild which demonstrates risk whilst minimising consequence (not a PoC, but an actual exploit) if that results in action to address the risk would actually be preferable to the ineffective pleading alternative which results in greater harms.
Reality is the domain of pragmatic choices and hard trade-offs.
The notion of regulations being written in blood is well-established. I'd strongly recommend the work of Charles Perrow, Normal Accidents and The Next Catastrophe in particular.
Further:
https://duckduckgo.com/?q=regulations+written+in+blood&ia=we...
https://avotraining.wordpress.com/2012/05/09/quite-often-the...
https://sharaevans.com/history-of-safety-regulation-written-...
https://www.maritime-executive.com/article/Safety-Regulation...
https://www.pilotsofamerica.com/community/threads/regulation...
https://aviation.stackexchange.com/questions/13081/why-do-pe...
https://books.google.com/books?id=vDnbtIzpnccC&pg=PA66&dq=re...
Maybe a randomized algorithm for ECC. Every so often changes how the ECC is computed?
The region nearby where privileged information is stored could have a speed limit on multiple sequential writes?
Add blocks that are more sucetible for those attacks to do early detection. A honeypot bit.
A DRAM-savvy OS might allocate memory so as to greatly reduce "critical data right next to untrusted data" situations.
CPU's which noticed write-heavy program behavior (and alerted the OS) would be interesting...but foolproof implementation could prove difficult.
(Ignoring the major drawback of not being able to replace/upgrade the RAM and CPU independently.)
Perhaps I should start snapping them all up … because market demand.
https://overclock-then-game.com/index.php/benchmarks/1-x5660...
As I understand it, it's just flipping random bits, wouldn't that have a very low chance of doing anything useful?
a) you can test and map out the memory layout, so its not entirely random
b) Often the system state can be manipulated in "safe" ways to increase the attack surface. E.g. you can cause the OS to fill lots of memory with a vulnerable data structure, like a page mapping, and at the same time reserve a lot of the memory (outside the area you want to target) yourself. This gives you a good chance something vulnerable will end up where you want it.
Afaik the Project Zero blogpost or presentation about it had some detailed examples.
It's really amazing to me that people can understand how a computer works to the level to be able to make that a viable attack. Very impressive.
If you don't need speed but need security...
They should do another set of tests on DDR5 and compare to DDR4.
Could you elaborate on what you mean by DDR5 "suffering" from SPECTRE? I believe that vulnerability is memory technology agnostic and theoretically would work if you could boot a CPU absent any memory at all, as the leaks occur through CPU caches.
"It's broken! Buy more of the broken stuff!"