If a modern CPU did no branch prediction, how much slower would it be? Just curious to know a ballpark figure, like 5x slower or 100x slower?
If a modern CPU did no branch prediction, how much slower would it be? Just curious to know a ballpark figure, like 5x slower or 100x slower?
If you designed a CPU to lack a branch predictor, you would certainly make different choices that would reduce the penalty of that, but there would still be a significant penalty.
Modern CPUs are all pipelined[2] because it allows a sequence of instructions to be executed (to some degree) in parallel. With out-of-order CPUs, the amount of parallelism that is possible for a sequence of instructions increases even more.
Without a branch predictor, you would have to stall the entire pipeline as soon as you see a branch instruction, until that instruction is finished. Obviously if you don't have a branch predictor, you're going to choose to have a shorter pipeline, but the penalty of each branch instruction will still be significant.
There's a lot of nuance to any proper answer to this question, and it's been years since I learned about this or thought deeply about this, so I'm probably not the right person to provide a deeper answer at this point. The performance impact would be significant.
[0]: https://news.ycombinator.com/item?id=34202230
[1]: https://stackoverflow.com/questions/11227809/why-is-processi...
So, there's no work required to "undo" those mistakes, but starting from an empty pipeline still means you're a dozen (or two) clock cycles from getting back to where you would have been with a successful branch prediction, which is why it is important for the CPU to predict correctly as often as it can.
Particularly the Pentium 4 suffered due to a long pipeline hence long stall. The successor to the Pentium 4 was an evolved Pentium III core with a much improved branch predictor and larger cache[1] which helped it outperform the Pentium 4[2].
[1]: https://en.wikipedia.org/wiki/Pentium_M
[2]: https://en.wikipedia.org/wiki/Intel_Core_(microarchitecture)
If you can execute ahead, through branches, you can resolve cache misses while the pipeline catches up, effectively reducing the cache miss penalty. If instead you have to stop at a branch, you're always gonna pay the full cache miss penalty, which can be significant.
Which, again, may not be an issue given your constraints.
However, even very dumb branch prediction tends to be well worth it for the vast majority of code.
https://arxiv.org/abs/2007.15919
This paper adds (more or less) a really big branch delay slot to fill the cpu with work while it's waiting for the branch to resolve.
Impossibility of Spectre-like attacks is a neat side-effect.
I wish they had included the geometric mean of their benchmarks, but I didn't see it anywhere, and I'm not going to run the numbers right now. Even if the speedup CFS offered on a 5-stage pipeline is "only" 25%, that is still huge... and on a larger pipeline, that delta would grow. "Not much" is drastically different from my interpretation of those results.
I do think security is extremely important, but I'm not convinced that things are currently so terrible that this is the only way forward, as the authors seemed to imply.
OTOH, I would enjoy seeing a return of an Itanium-style ISA that moves a lot of speculation from the hardware to the compiler. I think compilers are in a much better place now than they were when Itanium hit the scene, which did not help Itanium's problems.
Those applications tend to work, though, on the basis that either the compiler is generating fat binaries to support multiple architecture versions (e.g. Cuda), or some sort of IR, or compilation happens at runtime (e.g. OpenCL). It doesn't really work if you want to generate single binaries that will work performantly on a wide range of hardware versions - particularly important for users answering "how will application X work on future hardware Y", which really gets in the way of general-purpose use.
That's really the great advantage of putting more smarts in the hardware - you can evolve the processor design (often to improve performance) while executing the same binaries.
In my defense, 1.5X is "not much" when compared with 5 to 100X :-)
- there is a branch instruction every 10 instructions or so
- if you can't predict, assume you guess wrong half of the time
So every 20 instructions or so, you take the wrong branch. Sounds like it could be a significant overhead, from 10% - 50% depending on what the code inside the loop is doing.
For instance, if your code is doing only arithmetic operations, no load/store, such as computing factorial or other silly integer math, 20 instructions would probably take 20 cycles or less to execute, but the branch failure would cause a similar delay. So that's 50% overhead. In the other case, assume the 20 instructions are fairly slow (lots of memory accesses), then your branch penalty won't feel so bad.
In practice, branches aren't every instruction, so the impact will be less.
Modern pipeline lengths aren't documented publicly as often as they could be, but Zen 1 was a 19 stage pipeline, and it performed great for the time. But, I think Pentium 4's full pipeline was 31 stages. Regardless, I think CPU designers these days have better tools and techniques for dealing with long pipelines.
Not sure, because in that case the wrongly-predicting CPU is doing lots of wasted work, which contributes to eg waste heat. A non-predicting CPU could do more real work, than a wrong-predicting one?