Errata prompt Intel to disable TSX in Haswell, early Broadwell CPUs
techreport.com
techreport.com
I first heard about transactional memory when Sun had plans to implement it for its UltraSPARC Rock processor. There is a decent overview of the concept at http://en.wikipedia.org/wiki/Transactional_memory
The obvious initial targets for TSX optimization are server-class applications like transactional database servers.
Looking at your wiki links though, I would think this would be awesome for almost any high-performance multi-threaded program. Am I wrong about that?1. https://software.intel.com/en-us/blogs/2013/06/23/tsx-fallba...
1. http://natsys-lab.blogspot.ru/2013/11/studying-intel-tsx-per...
While the API of TSX (_xbegin(), _xabort(), _xend()) gives the appearance of implementing transactional memory, it is really an optimized fast path - the abort rate will determine performance. The technical term for what TSX actually implements is 'Hardware Lock Elision'.
If you are going to use TSX, don't use it directly unless you are writing your own primitives (i.e. transactional condition variables [0]), prefer using Andi Kleene's TSX enabled glibc [1] or the speculative_spin_mutex [2] in TBB.
[0] - http://transact09.cs.washington.edu/9_paper.pdf
[1] - https://github.com/andikleen/glibc
[2] - http://www.threadingbuildingblocks.org/docs/doxygen/a00248.h...
Mostly I wish there was a better fallback than locking when transactions fail, but I suppose this is more in line with thinking of them as optimistic locks (such as what I assume the speculative spin lock does).
A full-blown (say) lock-free hash table is almost insurmountable - I recall seeing an efficient reusable implementation a couple of years ago, but most before that had serious performance or correctness issues. However, if you take this into account early enough in the design of the system, you can usually do with much simpler lock free data structures (seqlock and friends).
transactional memory is, of course, a useful abstraction when you can't (or didn't) take this into account upfront.
Transactional memory is an access abstraction that presents memory as something over which transactions can be made. It says nothing about whether the implementation is lock-free (though the better implementations approach, or are lock-free). There exist software TM algorithms which lock during the transaction, which lock only briefly at the end of the transaction, and which are lock-free, and some which are a combination of these.
TSX (specifically HLE) only takes a lock if the transaction fails without taking a lock, in order to guarantee forward progress in the absence of an ill-behaved transaction (assuming fair locks). Were the software fallback implemented using lock-free techniques (possible with the more general RTM portion of TSX), this guarantee would extend to ill-behaved transactions as well.
Lock freedom is about progress, but in many practical cases (short time between load-linked/store-conditional pair, for example) it beats locks and everything else in performance -- basically, in the best case you do the same work as locks (i.e. synchronize caches among CPUs), and in the worst case, you spin instead of suspending the process.
Of course, if you do a lot of work between the load and store (or whatever pair of operations need to surround your concurrent operation), and there's a good chance of contention, then ... yes, it will not perform well. But that's not how you should use it (or how it is used most of the time).
And .. transactional memory is an abstraction the leaks differently than locks, but you must still be aware of the failure modes. It's higher level, but does not magically resolve contention.
So N years more till we see software transactional memory with dedicated hardware support.
Note that people have totally confused the terms. Originally "software transactional memory" meant using primitives like [strong] LL/SC to get lock-free, wait-free algorithms. Strong LL/SC has been proven to be a universal primitive which can be used to implement most known lock-free, wait-free algorithms.
"Hardware transactional memory" gave you access to lock-free, wait-free algorithms in a much more convenient manner. But really it's a difference of degree because strong LL/SC requires _significant_ hardware support.
These days people (like the PyPy project and most STM libraries) use "software transactional memory" for implementations which just use a mutex under the hood. It's emulated STM; they're just being fanciful.
I believe TSX is more like strong LL/SC. Maybe stronger. OTOH I'm not sure if it's universal enough to implement wait-free algorithms.
Gil Tene https://groups.google.com/forum/#!searchin/mechanical-sympat...
Cliff Click http://www.azulsystems.com/blog/cliff/2009-02-25-and-now-som...
It can be enlightening to hear from a real application of these technologies. I am sad to see it disabled. Had it in the back of my head to play around with it some time.
Really a non-statement statement. The only people who are using CPU instructions in their code will be low level maintainers, ie. kernel devs and other OS devs. So this doesn't effect really anyone except those folks (and of course the benefits it may have provided up the stack via layers of abstractions).
https://www.mikeperham.com/2013/12/31/rubys-gil-and-transact...
Say I bought a TSX enabled CPU specifically for that feature, I wonder if Intel will give me my money back... (they can have their broken CPU of course too)
Even if not, the US stick their nose anywhere, even where it does not belong...
Given that TSX is one of the features that distinguishes some of the more expensive Haswell SKU's, is Intel going to issue a refund for affected customers?
I see two about TSX:
HSD87 X No Fix Intel® TSX Instructions May Cause Unpredictable System behavior
Problem: Under certain system conditions, Intel TSX (Transactional Synchronization Extensions) instructions may result in unpredictable system behavior.
Implication: Due to this erratum, use of Intel TSX may result in unpredictable behavior.
Workaround: It is possible for the BIOS to contain a workaround for this erratum.
Status: For the steppings affected, see the Summary Table of Changes
HSD114 X No Fix Intel® TSX Instructions May Cause Unpredictable System behavior
Problem: Under a complex set of internal timing conditions and system events, software using the Intel TSX (Transactional Synchronization Extensions) instructions may observe unpredictable system behavior.
Implication: This erratum may result in unpredictable system behavior. Intel has not observed this erratum with any commercially available system.
Workaround: It is possible for the BIOS to contain a workaround for this erratum.
Status: For the steppings affected, see the Summary Table of Changes
HSD114 above seems to be the bug from the techreport article.Or is this such a non-issue that nobody cares?
The only real impact is that code that relied on TSX would need to rely on fallback methods of accomplishing the same tasks. Since there's probably not much of that floating around, there's very little impact at this time.
Is it even advantageous for an attacker to have the capability of issuing microcode updates to a target computer? What sort of attacks could you mount via microcode updates?
I'm guessing the private RSA keys are extremely well guarded, probably stored in a HSM that only allows signing, not key extraction, so that the keys cannot even be revealed to Intel.
On the other hand, who knows. There's certainly been cases of code signing keys on the loose (Adobe, etc) and even a compromised HSM host (Fedora, someone managed to sign compromised openssh .rpms)
It's plausible that a rogue microcode update could be used to bypass TPM static-root-of-trust protections, though, and a microcode update could certainly bypass TXT's dynamic root of trust. This might enable a bootable USB stick that would load malicious microcode and then reboot warmly enough to preserve the microcode and then launch the OS with a TXT bypass.
Even so, I don't really see the point. So far, essentially every BIOS can be freely (or freely using an exploit) reflashed from kernel mode, and a new image could contain malicious SMM code, and SMM code can also bypass both static and dynamic roots of trust.
Tamper with AES and randomness instructions.
Plant a very obscure privilege escalation exploit. Given the prevalence of java, activex and google nacl, escaping sandboxes is a big thing.
Tamper with the MMU to make a certain software invisible.
If they had the opportunity, it's not unlikely they did something like this. Perhaps just in a directed attack. They've done extensive firmware patching in the past in you believe last year's leaks.
A more interesting question is whether they didn't have to, because they had a say in making the silicon in the first place. The risk is there, it's not crazy to see it. Even FreeBSD who was the last mainstream OS to use the randomness instructions unadultered doesn't do that anymore.
Modulo the microcode signature, flash is better locked down these days: Flash updates are typically arbitrated by the firmware (though that's also just one signature away), requiring a reboot, while microcode updates are still free for all (in ring0 - realtek already lost a driver signing key once, why not again?).
[1] http://www.syscan.org/index.php/download/get/6e597f6067493dd... [2] http://mjg59.dreamwidth.org/30773.html
So yes, I'm quite aware of the immense set of faults in UEFI implementations (some of which are encouraged by UEFI's design, where more layers of UEFI are added to mitigate them).
But as an attacker I wouldn't want to assume that I run into any single of the many UEFI implementation quirks and adapt my attack to everyone of them.
And I really hope for Intel that Tianocore won't become an endless stream of portable UEFI security issues - otherwise the IBVs might get second thoughts about standardizing on a single codebase.
I would guess that the first layer of protection comes from limiting the scope of microcode updates. Perhaps there's a semiconductor engineer out there who knows more?
But a trojanized microcode update file inside an otherwise regular BIOS would be a nice hiding spot, hard to detect and analyze, at least for anyone outside Intel.
If you're an OS, then you need some kind of exploit to update SMM code. But if you're the BIOS, then you have complete control over what happens in SMM mode.
http://www.intel.com/content/dam/www/public/us/en/documents/...
It was last updated in June, so I guess it doesn't contain this latest erratum. Can't wait until it's updated... though I don't know if Intel is likely to disclose details.
Yeah, that's nice and descriptive. Damn. :/
Just reading through these issues, I see a handful that deal with specifically with external (proprietary?) hardware integrations, like with integrated Intel graphics boards. Why is the CPU loaded up extra jazz for with specific integrations like this? Is this to support a system-on-a-chip optional kind of configuration, or is this Intel trying to give themselves some sort of advantage in the graphics adapter market?
If it is the latter, then it serves them right to have these errata and instability due to adding all that extra competitive advantage nonsense on the chip.
TSX seemed like a once in a decade step forward, though as I understand it the restrictions with the cache size (and thus the amount of memory you can write to before the transaction gets too big) meant it wasn't very practical for much beyond optimistic lock acquisition.
For example, PyPy isn't planning on doing a TSX port even with their enthusiasm for transactional memory.
Also influencing my decision, I bought a 2600k a few years back, but never bothered overclocking it, which I guess was an admission that the excitement I found for hardware when I was a child was dead. I guess you either have the money, or the time, but rarely both.
It's disappointing that this microcode update isn't being done in such a way that you can re-enable it after agreeing to a disclaimer that it's not for production use. I'm not sure what the mechanism for this would look like, but given that Intel sold cards that unlocked Hyperthreading, I'm sure it's possible.
http://www.engadget.com/2010/09/18/intel-wants-to-charge-50-...
Edit: The article has been updated saying that it will be possible to enable TSX for development purposes on Haswell-EP at least.
Do you have more information about their reasoning behind this? From my point of view this is the highest profile software project to potentially make use of HTM, and I recall reading that the plan was to eventually introduce hardware acceleration.
http://pypy.org/tmdonate.html (Search for "haswell")
http://grokbase.com/t/python/pypy-dev/13bvt3kg70/pluggable-h...
It seems to boil down to:
* The cache size (which determines the amount of memory you can write to in a transaction before having to commit back) is insufficient, causing excessive transaction aborts.
* There is no mechanism to bypass the HTM, writing to memory within a transaction that is not rolled back. This exacerbates the small cache size, since all memory writes have a cost, not just the ones you want rolled back in the case of a transaction abort.
Interestingly, this does not bode well for HTM on a platform with many smaller cores, say a hypothetical 64 core ARM. Each core will have a tiny amount of L1 cache, severely limiting transaction size.
And many smaller cores is exactly where you'd want the benefits of HTM, since the overhead of synchronization is higher in proportion to the work each core can do.
First revision: MMX. It reused the same registers as the older x87 floating point coprocessor (even though the x87 transistors lived on the same die). As a result, legacy x87 code and MMX code had to transition using an expensive EMMS instruction.
Second revision: (well, ignoring some small changes to MMX) ... SSE. Finally got its own registers, but lacked a lot of real-world capability.
Third revision: SSE2, finally got to a level of parity with competing vector extensions (see, for example, PowerPC's Altivec).
And so forth.
I guess the take-home lesson for me is that these new TSX instructions are indeed fascinating to play around with, but I wouldn't expect it to blow the doors off. Intel will incrementally refine it.
(The incremental approach also gives Intel a chance to study how it's being used and keeps AMD playing catch-up.)
AMD's 3DNow had single precision floating point support, so it was actually somewhat useful. SSE followed 3DNow and added single precision support (as well as fixing the register stuff). SSE2 added double precision support.
Today, no one would use MMX instructions (since SSE is vastly superior). I expect Intel will continue to add TSX capabilities which will eventually produce some nice results for parallel code.
Yes, really:
http://ark.intel.com/products/80807/Intel-Core-i7-4790K-Proc...
...which is why I specifically bought it. So yes, I too am a little annoyed by this since I bought it specifically to develop TSX applications.
Hopefully you can upgrade to a Broadwell or later once Intel starts shipping fixed silicon. Haswells will be updated to the new microcode once a replacement is available. (At least, that's our plan.)
But yes, I'll be looking forward to the replacement Broadwell. My previous workstation was a Core 2 DUO E8400 which I just replaced with the 4790K.
Install the microcode by running: sudo apt-get install intel-microcode
Alternatively install a BIOS update once available.
Are there any scenario where Transactional Memory are useful in consumer environment?