I don’t know why, but there’s something really surprising to me about how far reaching the issues are, when you pull them back to first principles.
I don’t know why, but there’s something really surprising to me about how far reaching the issues are, when you pull them back to first principles.
A big group that turned it into "Intel fail" was due to various errors on Intel part, like moving verification of memory accesses to instruction retirement and similar things that even to uneducated person like me look like "tricks to make single-core performance shine".
I know a significant number of people who have disabled spectre and meltdown protections in favor of higher performance because they are not concerned at all about side channel attacks.
Single tenant workloads that run no untrusted code that are not public facing are not exactly a unicorn.
What Intel did was put all memory access auth checks to retirement, so appropriately crafted code could bypass protection bits in TLB, not just snoop elements of cache (AMD, ARM, POWER)
Intel's implementation of checking permissions in parallel with speculative loads was vulnerable because the transient effects could be observed through cache timing information.
All of these attacks, meltdown included, use "timing side channels from speculative execution."
POWER9 had the same problem (and therefore was similarly vulnerable to meltdown) for memory accesses that hit in the L1. i.e. userspace accesses to kernel addresses cached in the L1 could produce observable data-dependent effects because the data was available for further speculative execution before the permission exception was raised. This is why the Linux kernel put in place a mitigation for meltdown on POWER CPUs that involves flushing the L1-D cache on transitions to/from kernel space ("RFI flush").
To be clear, the vulnerability that Intel processors have because they "move memory access verification to instruction retirement" is called "meltdown." When I said that meltdown was present on IBM and some Arm designs, I meant that these designs are vulnerable to the same exploit because they have comparable problems with regard to memory permission checks. I was not referring to the various other spectre-type vulnerabilities, which are even more widespread.
So to summarize:
Ok, but "Intel fail" indicates meltdown was specific to Intel processors, when it was also present of IBM and some Arm designs.
That doesn't answer whether they were less optimized because someone on the Red team realized that was a good idea.
We can also talk about the Store-to-Load Forwarding involving a partially check physical address, exploited by the Fallout attack, which is an optimization patented by Intel since 2006 [1]. All these Intel optimizations are old, and sometimes patented (which reduces the chance that IBM, ARM or AMD will do the same) and are totally valid and safe as long as it's assumed that it is impossible to recover data used during the speculative execution.
The recovery of this data becomes possible only when the side channel Flush+Reload is discovered in 2014 [2] (which originally didn't target speculative executions). It is only 3 years later that Meltdown attack use Flush+Reload as a covert channel to exfiltrate data during speculative execution.
It seems to me that this timeline shows that it was not possible in 2006 to anticipate this issue. So it's not just a lack of education.
Perhaps a lack of security research at Intel to continuously challenge their optimizations?
But please note that there is no theory to simply determine whether these optimizations are a good idea or not. Criticism is easier afterwards when you know how the attacks are produced. But in 2017 this type of attack was something new.
That said, some of Intel's responses to the flaws (partial correction of vulnerabilities, lack of dialogue with researchers, etc.) are very worrying and I think these criticisms are much more legitimate.
I will admit that I'm not an expert in microarchitecture design, so to me the expected flaw would be "we do something wrong in speculative execution and end up bypassing the checks in destructive way", not through side channels. Didn't know about patents involved. And of course timing attacks make everything worse for all of us :/