BasicBlocker: ISA Redesign to Make Spectre-Immune CPUs Faster (2021)
arxiv.org
arxiv.org
Imagine being able to simultaneously visit 4 retail stores and dynamically select items depending on availability and pricing, arriving back home having spent the amount of time it takes to shop at 1.25 stores while burning 1.5x the fuel of a one-store trip.
There is no amount of ISA redesign or recompilation that can accommodate the dynamics of real-world trends in the same ways that speculative execution can. Instead of trying to replace speculative execution, I think we should try to put it into a more secure domain where it can run free and be "dangerous" without actually being allowed to cause trouble outside the intended scope. Perhaps I am asking for superpositioned cake here. Is there a fundamental reason we cannot make speculative execution secure?
What a fantastic analogy.
Any memory access leads to a time channel, which might be observable or not. As example, hyperthreading is known to create observable side channels even in L1 and L2 cache, since those are shared.
L3 cache is also shared between CPU cores on the same socket, so unless you can ensure the L3 cache data can never be shared you can not entirely eliminate this time channel.
Now getting back to speculative execution: All possible execution sequences must ensure to satisfy all possible cross-interaction rules to not make any time channel visible, which 1. includes restoring previous state fully and 2. not leaking any timing behavior. Just think in your mind of all possible cases, which would need to be verified (the complete instruction set) and if you can think of any sane time travel time leaking cache-aware separation logic to do this.
On top of that it has already been shown that fundamentally the hardware guarantees on cache behavior are broken (they are merely a hints).
Does this refer to some theorem? I'm very interested in how to model & prove formally that a given operation on a given system doesn't have timing side channels
You've said it yourself
> The fact that this also prefetches memory for us is amazing
To be secure, speculatively executed instructions that don't retire, have to have no observable effects, including those observable through timing. They cannot be allowed to modify the cache hierarchy in any way.
As soon as you allow any form of interaction, you're potentially leaking information through the timing side-channel.
Well, everyone patched things pretty aggressively, so exploitation isn't really practical. At minimum, that's why. Also, vuln research can take a long time to turn into exploits - there are so many existing primitives that people are exploring right now (ebpf, io_uring) that have practical exploitation primitives already designed, so there isn't much pressure to go for something that's already patched and that would require a lot of novel research to find reliable primitives for it.
As for the mitigations, just disable them? They're on by default because Linux doesn't know what your use case is, but if your use case isn't relevant to the mitigation's threat model please feel free to disable them. It is very simple to do so.
I also wonder to what degree we could partition between user mode and supervisor mode (and provide similar facilities to partition between user-mode and user-mode-sandbox, such as WebAssembly or other JITs), with the same premise. Let the kernel prefetch things but don't let userspace notice the speculated entries.
This is already sort of possible. The TLB flushing can take advantage of the PCID to determine, based on the process, whether the cache must be flushed - this provides process level isolation of the TLB.
I believe recent CPUs are increasing the size of some PCID related components since it's becoming increasingly important post-kPTI.
The only option I do see is something to prevent specific and limited memory as not being allowed into L3 cache altogether.
Since you're active and interested in this stuff: What is the state of art on cpusets for flexible task pinning on cores? "Note. There is a minor chance that a task forks during move and its child remains in the root cpuset." is mentioned in the suse docs https://documentation.suse.com/sle-rt/12-SP5/single-html/SLE..., but without any background explanation and I do not understand what stuff breaks on moving pinned Kernel tasks.
For example, a pre-fetch linked to a speculative execution would be tagged with a CPU-specific speculative execution identifier such that the pre-fetched data would only be accessible in that pipeline. If that speculative execution becomes realized then the tag would be updated (perhaps to zero?) to show it was actually executed and visible to all other CPUs and caches. In all other cases, the speculative execution is abandoned and the tagged cache cells become marked as available and undefined. Circuitry similar to register renaming could be used to handle tagging in the caches at the cost of effectively halving cache sizes.
In a more macro sense, imagine git branches that get merged back into main. The speculative execution only occurs on the branch. When the CPU realizes the prediction was good and doesn't need to rolled back, the branch is merged into the trunk and becomes visible to all other systems having access to the trunk.
It is secure in many, many contexts. For example, I have no concerns about speculative execution if I'm running a database or service, which is great since those are the areas where performance matters most.
Where it's troublesome is when you need isolation in the presence of arbitrary code execution. My suggestion is that if you ever find yourself in that scenario that you manage your cores manually - ensure that "attacker" code never crosses with anything sensitive on the same core. If you need that next level of security, enable the mitigations.
Pinning your cores is going to help a lot with the mitigations anyways - the TLB doesn't have to be flushed in circumstances where the same process on the same core is switched in, or something like that (someone please explain more, I forget and I don't want to look it up right now). There's some process context id cache blah blah blah, the point is that you can improve things if you pin to a core.
I think basically you're right. Instead of removing an amazing optimization let's find the areas where we can enable it, understand the threat model where it's relevant, and find ways to either reduce the cost of mitigations or otherwise mitigate in a way that's free.
I know I theoretically have the ability to set affinity, but have never explored it. Are there performance gotchas (outside of losing all core access)? Perhaps performance could even see minor improvement?
> Are there performance gotchas (outside of losing all core access)?
Depends on the workload and how the program is built. If you need to run tons of different processes concurrently (the 'user desktop' use case) it is not great. If your concurrency model requires creating hundreds of threads (Java, for most of the time it existed) it may not work well, especially if the threads don't have an even distribution of work (so you may see increased tail latencies). And, of course, if you have 8 cores but you isolate a workload to 6, you've just lost 2 cores that could have been used.
> Perhaps performance could even see minor improvement?
Considerable improvement, potentially. Less CPU synchronization means fewer L3 flushes, you'll get fewer context switches since your scheduler won't need to try to schedule other threads onto your core, fewer TLB flushes (relevant for kPTI/Meltdown mitigations), etc. If you lean into this approach you can get best in class throughput.
All in all, seems like a good idea, when can we staple this onto existing processors?
When you get out of abstract logical domains into real world physics, than the notion of machine state becomes fuzzier at higher clock speeds.
Good luck. =)