Maybe you meant out of order?
Maybe you meant out of order?
I think that encoding the parallelism into the instructions would generally not be considered as superscalar, it's an entirely different technique. Most texts I've read would explicitly separate out VLIW and EPIC from superscalar processors. There are plenty of ways of getting parallelism - for example one could probably view SIMD or even some CISC instructions as a form of this - as far as I know, VLIW and superscalar processors are just considered different from this perspective.
No. They use brute force, they have as many state machines as their superscalarness, and these SMs working simultaneous, saving results in OoO buffer, and periodically synchronize via special algorithm.
Each SM have full set of shadow registers and state of CPU.
To be precise, exists examples of non-orthogonal or "partially" OoO machines - Pentium MMX and IBM POWER. In them, CPU core divided to few parts, MMX parts dividing registers file, POWER have separate pipeline for memory operations, and these parts could do only in their limited space, but other mechanics is same as ideal orthogonal OoO design.
For example, when happen branch and unknown which side will run and have spare resources, SS running each sides (in separate OoO pipeline), and when will know finally effective side, results of other side just discarding, and results of effective side merging with other calculated data.
If happen branch, but resources limited, works predictor, which learn on previous runs and usually could predict with >80% accuracy, which side will run, so most probable side chosen.
When just happen spare resources on run, got last updated PC+1 and run other SM in parallel from there, than on some place (mostly on buffers fill), all data from OoO exec merging with main flow.
All this mean, carefully crafted program could see difference of OoO vs non-OoO execution (not exactly in memory, but by measure jitter, as merge OoO pipeline takes some time), but if not care, mostly have excellent compatibility and if fortunate, OoO will run as many times faster as number of OoO pipelines, which could be significantly large, ie in Pentium-3/4 was 4 or more pipelines.
And as side effects, this all gives high pressure on cache subsystem, and sometimes optimized in unsafe way (but cheap), so happen vulnerabilities, like Meltdown and Spectre.
It very much is superscalar. OoO Superscalar is what you described.
> The superscalar technique is traditionally associated with several identifying characteristics (within a given CPU): > Instructions are issued from a sequential instruction stream > The CPU dynamically checks for data dependencies between instructions at run time (versus software checking at compile time) > The CPU can execute multiple instructions per clock cycle
However above that is the simple definition, which is that as long as it executes more than a single instruction per clock it's superscalar. Even SIMD is taken as an example of a superscalar CPU.
As for the original claim the ia64 is very much a superscalar CPU. All VLIW designs are superscalar. VLIW can be also thought as OoO superscalar CPU with the reorder buffer and dependency analyzer ripped off and exposing the execution units explicitly in their gory details. Or alternatively we can also say that VLIW already won, we just added a big honking chip on top to JIT compile machine code into VLIW micro-ops.
Even in order non-superscalar cpus need some kind of dynamic checking if they allow out of order completion by allowing succesive low latency instructions to execute under the shadow of a preceding high latency instruction [1]. I'm not an cpu architect but this tracking is much simpler than the register renaming of OoO and only relies on hardware interlocks.
I think the grandparent is right in distinguishing VLIW, especially exposed pipeline ones, from superscalar as they have no tracking at all and just naively issue bundles; I think it is an useful distinction.
EPIC is more complicated as while it allows expressing intrabundle parallelism, it also allow dependencies and so it does need hardware interlocks. You could argue either way.
I think SIMD by itself should not be considered superscalar as it is still executing a single instruction is a single execution unit (compare the term superscalar itself to vector computation). Of course a superscalar CPU could have the capability to issue distinct SIMD instructions in parallel (for example larrabee).
CPU microarchitecture details are fun!
[1] there are many examples, starting with the CDC 6600; quake was also famous for implementing texture perspective correction by scheduling a division every 16 pixels and using fast approximation for the remaining pixels.
Yes, but ia64 was not pure VLIW design, it was EPIC - VLIW with additions. So compiler could suggest optimal path of execution, but if suggestion was not created, it using OoO (in reality things sophisticated, but on first look, something like this).
Other wide known VLIWs are nearly all ultraspecialized pure designs, without any OoO, because EPIC is very expensive in means of transistor number and cost of large die.
Example, AMD TeraScale GPU architecture, Radeon HD 2000 series, Radeon HD 3000 series, Radeon HD 4000 series, Radeon HD 5000 Series, Radeon HD 6900 series. https://en.wikipedia.org/wiki/TeraScale_(microarchitecture).