Intel to disable TSX by default on more CPUs with new microcode
phoronix.com
phoronix.com
(Not a perfect analogy I admit - opting out of a CPU security mitigation typically isn't going to risk anyone's life. At least not on most servers.)
> Intel's response to the FDIV bug has been cited as a case of the public relations impact of a problem eclipsing the practical impact of said problem on customers. While most users were unlikely to encounter the flaw in their day-to-day computing, the company's initial reaction to not replace chips unless customers could guarantee they were affected caused pushback from a vocal minority of industry experts. The subsequent publicity generated shook consumer confidence in the CPUs, and led to a demand for action even from people unlikely to be affected by the issue. Andrew Grove, Intel's CEO at the time was quoted in the Wall Street Journal as saying "I think the kernel of the issue we missed [...] was that we presumed to tell somebody what they should or shouldn't worry about, or should or shouldn't do".
(I'm not arguing equivalence between these two "post-consumer downgrades" -- just pointing out there's not only prior art for it occurring, but also prior art for repercussions/compensation.)
"According to different benchmarks, TSX/TSX-NI can provide around 40% faster applications execution in specific workloads, and 4–5 times more database transactions per second (TPS)"[1]
That sounds like a pretty big deal to turn off. Though it seems to be based on microbenchmarks. I wonder what the real world impact for a typical RDBMS is.
[1] https://en.wikipedia.org/wiki/Transactional_Synchronization_...
Curiously, postgres seems to perform okay (<10% drop) two years after the initial mitigations - perhaps due to newer versions of kernel-level fixes having less of a performance impact.
[edit] According to phoronix, the small penalty is mostly due to hardware-based mitigations in newer architectures.
[1] https://www.phoronix.com/scan.php?page=article&item=spectre-...
In other words it's been a total shit show. It doesn't have to be, but until Intel takes it seriously and produces a unflawed and secure realization it's going to continue to be a shit show. Even then you're still left with the Intel != AMD problem...
I think the first issue was that pthread locks using TSX behaved differently for incorrect usage, so that applications that had double-unlock bugs worked fine with the normal implementation but crashed with TSX glibc. These were app bugs, but a lot of people were unhappy that a glibc update caused their applications to crash, so some distros patched glibc to make time to fix the apps.
Then TSX was broken in Haswell and Intel disabled it, but it was a total mess, such that the bit in CPUID indicating support for TSX/RTM did not reliably indicate support for TSX/RTM (i.e. bit was set indicating support, but any TSX/RTM instruction resulted in SIGILL), so various versions of glibc ended up with various lists of CPU model numbers to determine when to avoid using TSX.
repeat for broadwell.
At this point more distros decided to turn off TSX in glibc entirely. But most did this by just removing the compile-time opt-in, which did not do anything for rwlocks. A couple distros noticed that rwlocks were still using TSX, and patched that too to remove lock elision entirely - but many did not, so rwlocks on haswell or early broadwell would still cause issues.
In my experience, a lot of the time RTM was not useful for performance because the transactions would be too large and abort. It was worse because in these cases glibc would try the transaction again several times before giving up, which could cause massive performance hits compared to just doing the normal lock to start with.
There were errata about TSX in skylake too, but I can't remember if Intel actually turned RTM off with microcode before now or just left it as is because no one was using it.
Eventually glibc was changed to always compile in support for TSX but require a runtime opt-in using an environment variable.
(this is from memory, the timeline is probably not quite right)
EDIT: I had a bit of deja-vu writing this, and it turns out I have whined about this on HN before lol: https://news.ycombinator.com/item?id=22694546
The issue is it's been YEARS now that intel claims it's ready then turns it off in chips shortly thereafter. It's basically vaporware at this point.
So I just wrote the more complex reader/writer lock system and went on my way. Turns out to have been the right answer...
At this point, TSX is like a new Google product launch - "This, too, shall pass." I don't know why anyone would bother spending much time with it after all these generations of "TSX! Wait, no... uh..."
It turns out that it doesn't really work. Surprise! not really
Their issue IMO was not really scoping down TSX... the project was far too ambitious.
https://lwn.net/Articles/534761/ http://halobates.de/adding-lock-elision-to-linux.pdf
And then try debugging some seemingly unrelated issue at the memory/instruction level, and drop the assumption that everything below your level of abstraction works perfectly as you expect.
And yes, this is very similar to mis-predicted branches and the speculative execution vulnerabilities. I have honestly tried to understand those at the instruction level, and felt like I did at one point, and then a few months later felt like I forgot some key details and couldn't put it together again... like I get the idea, but can't quite put together the assembly to demo the problem on my own, it's a bit of a mindfuck ...
That said, hardware support for atomic use of larger memory regions than is currently supported is something that is worthy as a problem to solve.
In a real database, you will run into transaction aborts much more frequently to say nothing of the correctness concerns raised by others in this thread.
[1]:https://web.archive.org/web/20161110144922/http://pcl.intel-... [2]:Improving In-Memory Database Index Performance
Cliff Click — The Azul Hardware Transactional Memory experience
https://www.youtube.com/watch?v=GEkeOHw87Sg
TL;DW: Transactional memory sort of promised lock elision to speed up gratuitously lock heavy code, particularly for read heavy workloads. Think a caching hash map that only has the occasional writer. It didn't live up to that since the max transaction size is really easy to overflow (far easier than you may think), and general code has a nasty habit of writing metrics even in read cases which means that transactions all write to the same address, which means constant transaction conflicts, which means slower code than just using a lock. Cliff Click then makes the argument that the right move in the vast majority of cases if you're having lock contention is to rewrite the algorithm so that threads (or at least cores) communicate via messages and share nothing rather than trying to use transactional memory so that you can cleanly fork out across a cluster rather than being constrained to a single box still after the rewrite. And for context he comes to that conclusion while working on massive 768 core boxes where the value add is "we sell you a giant non-NUMA single box that absolutely eats Java for breakfast, lunch, and dinner".
Even if you read the TL;DW, I still encourage you to listen to his talk; he's an incredibly smart engineer and I learn something pretty much every time I listen to him.
[0] The Sea of Nodes and the HotSpot JIT
[1] A JVM Does That?
[2] A Crash Course in Modern Hardware
[3] Bits of Advice for the VM Writer
[0] https://www.youtube.com/watch?v=9epgZ-e6DUU
[1] https://www.youtube.com/watch?v=-vizTDSz8NU
Preston Briggs came out of that group, and did his PhD thesis on register allocation via graph coloring, one of the early papers on that technique: https://www.cs.utexas.edu/users/mckinley/380C/lecs/briggs-th.... That's a great, very readable paper.
Cooper & Simpson's paper on SCC-based value numbering is one of the more readable works on the subject: https://www.cs.rice.edu/~keith/Promo/CRPC-TR95636.pdf.gz
Vikram Adve worked there. He was the PhD advisor to Chris Lattner, who developed LLVM.
And I heard that the HTM worked pretty well, but to be fair nobody really got a chance to prove them wrong. I'm sure Intel was pretty confident in their implementation too.
Do people even read past the headlines these days?
With Intels wave of TSX disabling patches going over all the previous CPU generations, AMD apparently got it right by never bringing their own version into production at all.
How would you get memory protection or debugger support then?
Or the very-likely use-case of rendering live (i.e. untrusted) web-content in-game? That would require a modern engine like Chromium which in-turn necessitates per-process isolation.
Also, no-one wants to have to reboot their entire computer system just because some application code crashed. Even if preemptive scheduling and the MMU works in ring 0 all processes in ring 0 can manipulate the MMU registers, so a misbehaving process can still bring down the entire system. Bad idea.
Even dedicated HPC applications need it - remember that the HPC/scientific-computing/batch-job-processing model is the whole reason that operating-systems exist today: they're highly evolved successors to job-control programs (e.g. Chippewa Operating System). So I don't think they'd want to give up all the improvements of the past 50+ years.
Ultimately we're just arguing about the principle of how the user who owns the computer should be in ultimate control over the hardware - but unfortunately Apple has demonstrated that principle is unnecessary to make obscene profits and also runs contrary to the realities of running a large secure platform today. We shall see how this ends-up...
The root user still can access read/write the whole memory of the system anyway via /dev/mem, load kernel modules, and do practically anything on the system. In that context, a change of privilege is useless in practice, and can lead to performance degradation.
This will lead performance benefits in most system daemons, notably programs that interact closely with the kernel like systemd.
From the article:
> A memory ordering issue is what is reportedly leading Intel to now deprecate TSX on various processors.
And from the Intel document linked in the article:
> The default RTM force-abort behavior can be optionally disabled by setting MSR bit TSX_FORCE_ABORT.SDV_ENABLE_RTM=1. However, when RTM force abort is disabled in this way, RTM usage may be subject to memory-ordering correctness issues. Due to these issues, this unsupported mode should not be enabled for production use.
But I doubt we'll see microcode patches that fix older CPUs. I have a Haswell where TSX was permanently disabled via microcode and s Comet Lake that straight up came from the foundry without TSX.
I will reconsider once a PoC comes out that affects either nginx or sshd.
I think you’d have to somehow disable this file in WinSxS as well as the installer/image files too.
[1] https://www.arm.com/why-arm/architecture/security-features/a...
>Transactional Synchronization Extensions (TSX), also called Transactional Synchronization Extensions New Instructions (TSX-NI), is an extension to the x86 instruction set architecture (ISA) that adds hardware transactional memory support, speeding up execution of multi-threaded software through lock elision.
How is that related to CCA, which seems to be a security feature?