https://arxiv.org/abs/2007.15919
This paper adds (more or less) a really big branch delay slot to fill the cpu with work while it's waiting for the branch to resolve.
Impossibility of Spectre-like attacks is a neat side-effect.
https://arxiv.org/abs/2007.15919
This paper adds (more or less) a really big branch delay slot to fill the cpu with work while it's waiting for the branch to resolve.
Impossibility of Spectre-like attacks is a neat side-effect.
I wish they had included the geometric mean of their benchmarks, but I didn't see it anywhere, and I'm not going to run the numbers right now. Even if the speedup CFS offered on a 5-stage pipeline is "only" 25%, that is still huge... and on a larger pipeline, that delta would grow. "Not much" is drastically different from my interpretation of those results.
I do think security is extremely important, but I'm not convinced that things are currently so terrible that this is the only way forward, as the authors seemed to imply.
OTOH, I would enjoy seeing a return of an Itanium-style ISA that moves a lot of speculation from the hardware to the compiler. I think compilers are in a much better place now than they were when Itanium hit the scene, which did not help Itanium's problems.
Those applications tend to work, though, on the basis that either the compiler is generating fat binaries to support multiple architecture versions (e.g. Cuda), or some sort of IR, or compilation happens at runtime (e.g. OpenCL). It doesn't really work if you want to generate single binaries that will work performantly on a wide range of hardware versions - particularly important for users answering "how will application X work on future hardware Y", which really gets in the way of general-purpose use.
That's really the great advantage of putting more smarts in the hardware - you can evolve the processor design (often to improve performance) while executing the same binaries.
In my defense, 1.5X is "not much" when compared with 5 to 100X :-)