The idea here is to essentially create a branch delay slot instruction, which then can be used to set the latency of a branch such that it doesn't require prediction to not stall the pipeline. Like so:
dec r1
basicblock 5 # next five instructions WILL execute, branches after that, if any
jz exit_loop
this
will
always
run
-- branch takes effect here
Then the compiler may commonly reorganize the code such that the variable branch delay instructions perform useful work.Architectures with ISA-specified delay has significant issues as you mention.
the measurements show that they can generate bb of about 10~20 instructions (optimistic numbers as they measure an average of 5) which allows them to move up the branches of about 10~20 instructions. As with this ISA the bb determines the bound of the instruction window, then the instruction window of this ISA is limited around 20~40 instructions (current bb plus next bb). But modern superscalar processors have a instruction window > 100 instructions to provide high performance.
The low performance loss of their model compared to the branch prediction model may perhaps be explained by the fact that they use an in-order CPU that makes very little use of the ILP (Instruction Level Parallelism).
Moreover, it only addresses the problem of speculative execution, but there are other types of transient execution.