When we talk about speculation and out of order execution we imagine the CPU is a kind of giant dependency-graph engine but this isn't actually true. The genius of modern hardware design (although kicked off by Tomasulo in the 60s) is that you can "compress" this idea into a real circuit with a finite number of gates and SRAM etc.
The quality of the branch prediction lets you get away with a lot on a fairly dumb "throw stuff on a buffer, speculate across memory accesses, pop (commit) off the buffer when the pie's finished cooking" model inside the processor.
The draw of this and originally Transmeta is that you have an extremely wide dumb processor in front of a smart software frontend that can do much more complicated work and scheduling based on the all-important runtime information that static VLIW sorely lacked (and thus lead to Intel doing all kinds of stuff with Itanium, hence it ending up EPIC rather than a true VLIW spiritually).
Now, software is slow, the way Transmeta got around this is by having a physical cache (the "Tcache") for storing translated instructions in.
https://www.cs.cornell.edu/courses/cs6120/2019fa/blog/transm...
https://www.realworldtech.com/crusoe-intro/5/ Some reverse engineering from the time. Note the amount of nops in the firmware.
https://safari.ethz.ch/digitaltechnik/spring2019/lib/exe/fet...