Or is the worry on the other side; that processors have gotten so out-of-order that only huge dedication to guesswork can keep the beast sated? I don’t see this as a million miles from software techniques in JiT compilers to optimistically optimize and later de-deoptimize when an assumption proves wrong.
I think you might be right to be nervous if you wrote programs that took fairly regular data and did fairly regular things to it. But, as Itanium learned the hard way, programs have much more dynamic, emergent and interesting behaviour than that!
Since processors are expensive and hard to change, they do tricks to allow themselves to be used more efficiently in common cases. That seems like a reasonable behavior to me.
However, it wasn't exactly a raging success, with I think the predicted amazing compiler tech not materialising, but maybe it is the right answer, but the implementation was wrong? I'm no CPU expert...
I do think a big part of the problem is that people want to distribute binaries that will run on a lot of CPUs that are physically really different inside. But nowadays there's JIT compilation even for JavaScript, so you could distribute something like LLVM, or even (ecch) JavaScript itself, and have the "compiler scheduling" happen at installation time or even at program start.
There have been a small number of attempts since Itanium, like NVIDIA's Denver, which make for much better baselines. I don't think those are anywhere close to optimal designs, or really that they tried hard enough to solve in-order issues at all, but they at least seem sane.
The interesting probabilities are all decided at runtime.
Now we have AI workloads there is a place for a big lump of dumb compute again, but not in general purpose code.
Can branch prediction be turned off on a compiler or application level? If you're optimizing for energy use that is. Disclaimer: I don't actually know if disabling branch prediction is more energy efficient.
The data cache memory is one of the solutions to avoid the extremely long latency of loading data from a DRAM memory.
The alternative to a data cache memory is to have a hierarchy of memories with different speeds, which are addressed explicitly.
The latter variant is sometimes chosen for embedded computers where determinism is more important than programmer convenience. However, for general-purpose computers this variant could be acceptable only if the hierarchy of memories would be managed automatically by a high-level language compiler.
It appears that writing a compiler that could handle the allocation of data into a heterogeneous set of memories and the transfers between them is a more difficult task than designing a CPU that becomes an order of magnitude more complex due to having a hierarchy a data cache memories and a long list of other hardware mechanisms that must be added due to the existence of the data cache memory.
Once it is decided that the CPU must have a data cache memory, a lot of other hardware design decisions follow from it.
Because there is an inverse relationship between the load latency and the data cache memory size, the cache memory must be split into a multi-level hierarchy of cache memories.
To reduce the number of cache misses, data cache prefetchers must be added, to speculatively fill the cache lines in advance of load requests.
Now, when a data cache exists, most loads have a small latency, but from time to time there still is a cache miss, when the latency is huge, long enough to execute hundreds of instructions.
There are 2 solutions to the problem of finding instructions to be executed during cache misses, instead of stalling the CPU: simultaneous multi-threading and out-of-order execution.
For explicitly addressed heterogeneous memories, neither of these 2 hardware mechanisms is needed, because independent instructions can be scheduled statically to overlap the memory transfers. With a data cache, this is not possible, because it cannot be predicted statically when cache misses will occur (mainly due to the activity of other execution threads, but even an if-then-else can prevent the static prediction of the cache state, unless additional load instructions are inserted by the compiler, to ensure that the cache state does not depend on the selected branch of the conditional statement; this does not work for external library functions or other execution threads).
With a data cache memory, one or both of SMT and OoOE must be implemented. If out-of-order execution is implemented, then the number of registers needed to avoid false dependencies between instructions becomes larger than it is convenient to encode in the instructions. so register renaming must also be implemented.
And so on.
In conclusion, to avoid the huge amount of resources needed by a CPU for guessing about the programs, the solution would be a high-level language compiler able to transparently allocate the data into a hierarchy of heterogeneous memories and schedule transfers between them when needed, like the compilers do now for register allocation, loading and storing.
Unfortunately nobody has succeeded to demonstrate a good compiler of this kind.
Moreover, the existing compilers have frequently difficulties in discovering the optimal allocation and transfer schedule for registers, which is a simpler problem.
Doing efficiently the same for a hierarchy of heterogeneous memories seems out-of-reach for the current compilers.
And last not least, unknown memory latency is not the only source of problems, branch (mis)predictions are another. And they have the same remedies as cache misses: multithreading and speculative execution.
So if you wanted to get rid of branch prediction as well, you could come up with something like the CRAY-1.
However, for this, fine-grained multi-threading is enough. Simultaneous multi-threading does not bring any advantage, because the thread with the mispredicted branch cannot progress.
Out-of-order execution cannot be used during branch mispredictions, so like I have said, both SMT and OoOE are techniques useful only when a data cache memory exists.
Any CPU with pipelined instruction execution needs a branch predictor and it needs to execute speculatively the instructions on the predicted path, in order to avoid the pipeline stalls caused by control dependencies between instructions. An instruction cache memory is also always needed for a CPU with pipelined instruction execution, to ensure that the instruction fetch rate is high enough.
Unlike simultaneous multi-threading, fine-grained multi-threading is useful in a CPU without a data cache memory, not only because it can hide the latencies of branch mispredictions, but also because it can hide the latencies of any long operations, like it is done in all GPUs.
Fine-grained multi-threading is significantly simpler to implement than simultaneous multi-threading.
Alloc/GC: https://github.com/dotnet/runtime/blob/main/docs/design/core...
Reg alloc: https://github.com/dotnet/runtime/blob/main/docs/design/core...