(The one exception is also interesting. AMD processors allow speculative reads past the end of x86 segments and past BOUND instructions, which of course no-one uses these days. This suggests there may have been a deliberate decision to block them in the more important cases.)
Somebody messed-up big time. Or from a business point of view did they? Intel current problems are manufacturing and the continuously lower power of their "legacy" processor (except due to manufacturing problems, this "legacy" is still mostly the current one) makes it so that people are buying more. Of course there is AMD back in the game, but the market demand is large enough; plus AMD would have been there anyway, and, in the fiction that Intel did take good parts of the perf hit upfront instead of the secu vulns, as competitive as in the current situation.
The people most annoyed are the users. Intel got away by pretending this was not really defects in their product but only new SW tricks that they will help defend against, and their clients just let them say that without much complaint (well, I guess big ones got some rebate...) but security researchers and/or processor designers know very well this is bullshit (see the vulns papers and FAQ) and that they simply fucked-up big time on Meltdown, MDS, etc. I don't care that a few other vendors did some of the same mistakes: they are still mistakes and design flaws, and not even something new.
Pretty much the only new shinny thing in this stream of vulns was Spectre and the few variants that appeared quite early on (but NOT Meltdown&co). The rest are design flaws that comes from the "oh not a big deal to leak that potentially privileged data, we will drop anything and trap before any derivative can go out anyway" mentality. Yeah, no, I'm sorry but the funding paper about speculation already told to not do that :/ Either they did not do their homework, or they voluntarily chose to violate the rule.
... if you value security over raw performance. Clearly Intel has decided at some point that it was worth playing with fire in order to get ahead in benchmarks. In their defence it seems to have worked reasonably well for them for quite a while.
>The people most annoyed are the users.
I wish, but I wonder how much of that is true. Are most users even aware of these problems? They get patched automatically by OS vendors and then most of the time they won't hear about them anymore. I think the "nobody gets fired for choosing Intel" will probably still prevail for quite some time.
They built a bunch of tech debt into their processors to boost their numbers, and now they hens are coming home to roost.
What I'm wondering is how many changes this will make to their product roadmap, and to what degree it will make next generation chips look lackluster compared to what people (think they) have now.
It seems to be down to the notoriously buggy TSX (hardware transactional memory) in Intel CPUs.
They had an additional mode that would transparently convert many spinlocks into transactions without code changes - that is now gone.
As core counts increase spinlocks and other synchronization primitives simply become too expensive. We'll need transactional hardware support eventually.
Scaling workloads does not require transactional memory and certainly doesn't require a vulnerable implementation of it. HTM might be the easiest way to scale a relatively naive algorithm, but the most scalable synchronization is none at all (or vanishingly infrequent) — and that works just fine with conventional locks and atomics (both locked instructions and memory model "atomics" such as release/acquire/seq_cst semantics).
TSX brings hardware supported optimistic locking and breaks the latency imposed by MESI and related protocols in use today. Of course its great if you can get away with no synchronization at all - but then you might as well just use a GPU. TSX helps with those non-trivially parallelized problems that are still best performed on a CPU.
Obviously many workloads require some coordination, but often something as trivial as allocating one of a given resource per CPU is sufficient to avoid most contention even on 100s of CPU core machines. Profile; improve. The same is required with HTM.
Regardless of your thoughts on HTM and scaling technology, TSX is broken from a security standpoint, which is the primary subject of the fine article. HTM != TSX.
I’m going to nitpick, though.
> I explicitly mentioned memory model atomics in addition to locked instructions in an attempt to prevent getting hung up on locked atomics. I guess that didn't work.
By “memory model atomics” do you mean atomic loads and stores rather than, say, compare-and-swap or atomic-increment? Because C++11 compare-and-swap and atomic-increment operations take a memory ordering parameter, yet they still generate the same `lock cmpxchg` and `lock inc` instructions.
But the issue isn’t locks; none of those instructions actually lock the entire memory bus like in the old days (...unless you pass an unaligned address!). They lock the cache line, which causes cache line thrashing if many processors do it at the same time, and that’s the biggest source of overhead. But plain stores also lock the cache line. A compare-and-swap is more expensive than a plain store, but not that much more.
Yet another Intel exclusive side channel vulnerability might help change that. This side channel stuff is terrible for cloud operators. Every time they have to adopt another layer of mitigation some fraction of their capacity disappears in a puff of shame and excuses.
The big migration will happen in about 2-4 years. Typical enterprise schedule is 3-5 year cycles, zen2 launch was 2019.