Are you referring to this?
Performance Details
Overhead
Naturally, protecting an indirect branch means that no
prediction can occur. This is intentional, as we are
“isolating” the prediction above to prevent its abuse.
Microbenchmarking on Intel x86 architectures shows that our
converted sequences are within cycles of an native indirect
branch (with branch prediction hardware explicitly
disabled).
For optimizing the performance of high-performance binaries,
a common existing technique is providing manual direct
branch hints. I.e., Comparing an indirect target with its
known likely target and instead using a direct branch when a
match is found.
Example of an indirect jump with manual prediction
cmp %r11, known_indirect_target
jne retpoline_r11_trampoline
jmp known_indirect_target
One example of an existing implementation for this type of
transform is profile guided optimization, which uses run
time information to emit equivalent direct branch hints.
-- https://support.google.com/faqs/answer/7625886
Note the parenthetical in the first paragraph. The way I read the above, they're merely saying that the performance is the same as disabling branch prediction for the section of code. Which would be really slow. But, presumably, faster overall than _actually_ disabling branch prediction as there's probably no good way to explicitly do that in as localized a manner.[1]
And as I interpret this, the reference to FDO is unfortunately confusing and perhaps better placed under the heading of correctness. They're merely saying that the transformation of indirect branches to direct branches is something compilers already do to steer speculative execution down the correct path. Except in the case of a retpoline it's steering speculative execution into a trap.
[1] I'm no expert in this area (nor do I do much assembly programming) but AFAIU branch predictors these days have substantial caches of their own. Disabling branch prediction or flushing the predictor's caches is probably much more costly than simply steering the predictor into a trap. The single pipeline bubble you create is better than the many bubbles that would be created if the branch predictor had to warm up again for later code.