You are correct that VLIW, e.g., as in Itanium,
flopped and out of order execution, speculative
execution, branch prediction, etc. won out.
The 9:1 speed up I reported was directly from Ebcioglu:
For a while he was in our AI group at Watson. I don't
recall just what he did, but he may have had an intermediate
step that, for a given program, analyzed the stream of
370 instructions and rewrote them for the 24 way
VLIW.
An idea I had him consider was very long addresses. Or,
who the heck really wants the addresses in main memory
to be 0, 1, 2, ...? Of course, no one! That's why
we have lots of work, from old link edit relocation to collection classes, memory management (garbage
collection), etc.
So, what do we actually do? Sure,
darned near everything comes out of level 1, 2, 3 or so
cache which works with just a hash of the
main memory address.
So, since we are going to hash the main memory
addresses anyway, let's just do that! So, have
a very long address, say, 1024 bytes long.
So, the 1024 bytes are immediate from the
source code! That is, just concatenate,
say, address space name, program name,
function name, class name, instance name,
member name, key name (as in key-value pairs).
That's the address. Then hash it.
Ebcioglu
checked how fast the execution logic could be,
and it seemed okay.
I know; I know; I left out
a lot and need a better explanation! I'm
concentrating on my software for my project
and have gotten away from such low level hardware
issues; heck, I don't even remember if
we hash the real or the virtual address!
But, the OP was a little light on a lot of
history maybe relevant to the work of
Soft Machines.