Interesting that AWS has been mitigating for side channel attacks since before they became a big news item. Curious about Azure and GCP's stance on this
Interesting that AWS has been mitigating for side channel attacks since before they became a big news item. Curious about Azure and GCP's stance on this
Or, maybe, they just thought the lack of deterministic performance created billing/accounting/customer service problems. (One hyperthread can just about completely starve in many circumstances).
You pick a VPS because you want to save costs and share cores.
So it’s not so much AWS choosing as you the customer choosing.
I ran the numbers back in 2015 and persuaded the previous employer to go for dedicated tenancy with all performance-critical and privacy-sensitive workloads. I was effectively hedging against unknown but practically guaranteed cross-VM attacks to pacify a paranoid regulator.
Then Rowhammer happened. Less than a day later, our contact with the regulator comes to me asking how it affects us. Being able to answer - with absolute confidence - that it did not, was one of the proudest moments of my career. And the turnaround from "regulator comes asking awkward questions" to "regulator is happy and sees no reason to ask again" of less than 48 hours must be some kind of record too.
"Are We Susceptible to Rowhammer? An End-to-End Methodology for Cloud Providers"
https://arxiv.org/pdf/2003.04498.pdf
Although their answer in this paper was diplomatic, my interpretation is that they confirm it as a problem. Their conclusion was it would not be as bad it was considered at the time. To be revisited on the context of this more recent work.
Edit: Adding main reference
"BLACKSMITH: Scalable Rowhammering in the Frequency Domain" https://comsec.ethz.ch/wp-content/files/blacksmith_sp22.pdf
The looming thread of side-channel attacks on SMT systems has been known since well, before we had SMT systems (because it also can apply to Co-Processor, and non SMT multi core systems).
The difference between back then and today is "we believe it's possible but haven't found a way yet" and "there are multiple known ways", as well as it being wide spread known instead of just in some communities.
The reason we still shipped the problematic CPU's is because improvement in perf. and as such competitiveness and revenue on the short term where more important.
There also was a shift in what people expect from security and which attack vectors are relevant. For example user applications a user installed where generally trusted as much as the user. While today we increasingly move to not trusting any applications even if they are installed by a trusted user and produced by a trusted third party. Similar running arbitrary untrusted native code from multiple untrusted users and "upholding" side-channel protection wasn't often an important requirement in the past.
I think the big difference is that we have attacks with high bandwidths of exfiltrated information-- instead of very slow, difficult to control leaks.