Hand-Optimizing VLIW Assembly Language as a Game
silverspaceship.com
silverspaceship.com
For example Linus Torvalds has made this assertion a few times that I've seen. However, it seems to me (at least by their own telling) that their focus was almost entirely on trying to do x86 better than x86 and doesn't really seem like a straight forward or easy task.
What I wonder about is VLIW doing its own thing (e.g. not trying to run Windows games) but running applications intentionally developed for it. Perhaps there are use cases where VLIW works very well... or at least well enough to a reasonable alternative value proposition.
VLIW blew up for obvious[0] reasons - memory access latency is a dynamic thing and can't be anticipated statically very well. Therefore an in-order VLIW chip with its basket of instructions (one of which being a mem access) will stall, the whole basket freezes, the subsequent baskets are held up in the queue, etc.
That's why OOO processors won, a memory miss can to a degree be worked around by executing instructions ahead in the queue, if they're not dependent on the stalled instruction.
[0] Blindingly obvious to me when it was pointed out, but it did need pointing out I have to admit.
Mill has no answer for the unpredictable cache misses. Their belt idea is cute and interesting but inconsequential and irrelevant.
The main challenge of the modern microprocessors is not in how many ALUs you run in parallel nor how you model and implement your register file, but how to keep the ALUs fed with data from the memory hierarchy. VLIW doesn't help you with that at all, which is why it is largely irrelevant and a tiny niche.
For applications like GPU, you address the feeding problem with multiple threads through SMT/SIMT and multiple cores, and SIMD.
For applications where a single thread performance is important, you use OOO and/or speculative execution to maximize the feeding of the data.
If you only care about running a particular loop and you only need to meet a fixed cycle budget like DSP application, VLIW may make sense when you want to save some engineering time and you have relatively low volume. Once volume is high enough, you will likely choose an in-order superscalar instead, for better performance or even simpler implementation, ease of programming, better backward compatibility and higher instruction density. VLIW has become more of a niche even in DSP. E.g. see https://www.eetimes.com/a-case-for-superscalar-dsps/ which is now almost 20 years old article.
The basic idea behind VLIW of moving as much work as you can to the compiler is not wrong, but VLIW turned out to be a bad way to divide work between the pipeline and the compiler.
OOO and speculative execution are very energy-intensive per unit of compute, and hard to implement securely. Even if a VLIW-ish cpu like the Mill is still ultimately bottlenecked by memory bandwidth, avoiding pretty much every known drawback of OOO is a huge deal.
As for the security aspect of the speculative execution, OOO outperforms VLIW dramatically for a single thread performance even without any speculative execution. Moreover, since the discovery of Spectre, there's been plenty of research into how to do speculative execution without leaking:
https://ieeexplore.ieee.org/abstract/document/8806916
https://ieeexplore.ieee.org/abstract/document/8574559
https://ieeexplore.ieee.org/abstract/document/9138957
https://dl.acm.org/doi/abs/10.1145/3316781.3317914
It's not yet clear what approach will actually win out and will be adopted, but it's clear there are multiple viable solutions such that this is a solvable problem. The whackamole of fixing / addressing individual spectre-like vulnerability likely will continue for a while - after all, processor folks didn't have that much time given the development cycle for new pipeline/architecture is measured in years - but it's unlikely that side-channel issue will make speculative execution not viable.
Edit: s/SIMT/SMT/
It would seem that instruction compression can be used to recover that density in many cases. Even Itanium did something like that IIRC, and of course the Mill is noticeably reliant on cheap representation of "unused" instructions.
> It wastes the die space - superscalar + SIMT are inherently more efficient when you have more than 1 thread
I don't believe this at all. SIMT encompasses approaches such as barrel processing, and wide-SIMD plus predication as an alternative to branching code, that while effective in some specialized cases are hardly "inherently" preferable in general.
Yes, they have. They have separated fetch and use operations and a completely rethought writing cycle.
Of course, I have no idea how that actually works on practice.
> Mill has nothing to show for it.
so far.
> Mill has no answer for the unpredictable cache misses
And they state so, quite explicitly (and as no architecture can ever, as I said in another comment). so you can hardly criticise them for that.
> Their belt idea is cute and interesting but inconsequential and irrelevant.
It's about avoiding register forwarding, so chopping out a major chunk of very high speed hardware (AIUI).
> but how to keep the ALUs fed with data from the memory hierarchy
No, it depends entirely on the job. If you have tight loops with lots of ILP then you win.
(as for the rest, IDK).
But that's the whole point of OOO isn't it? Keep executing stuff down the line while previous instructions are waiting for the cache to get filled.
There's been some newer OoO cores that'll speculate on the data contents of the load, but you'd be surprised how little that helps all things considered.
That view is only a very partial understanding of the bigger story, and is missing the key benefit of OOO.
The key benefit of OOO and speculative execution is that they allow you to generate multiple outstanding L2/L3 misses in flight. Memory latency is hundreds of cycles, and if you can get one more memory read in flight while there's already one miss, you cut the effective stall time for memory significantly. OOO and speculative execution help not only because they allow the pipeline to run more instructions while being stalled, they help because that extra execution can create multiple outstanding cache misses, effectively reducing the stall cycles due to cache misses.
DSPs, and of course GPUs, do exist and again you're right, for special purpose but important niches of course, but I was referring to general purpose CPUs but didn't make that clear. Thanks.
The biggest takeaway for me was the importance of scalability. As OOO CPUs get better, they can become increasingly parallel when running the exact same assembly. While VLIW like Itanium was limited to 3 ops/instruction, which is already behind today's OOO CPUs. (Although Mill's 33 ops/instruction sounds like it'd take much longer to beat.)
My takeaway was that NVidea's Volta looked on the right track, where it added 6-bit dependency bitmasks to instructions to make building the dependency graphs (primary cost to OOO) more efficient.
Itanium could group 3 instructions per 128 bit "bundle", but multiple bundles could be run in parallel, and the compiler had to insert explicit "stops" when that wasn't allowed. The architecture was designed specifically so that future processors could run old code with more parallelism.
https://en.wikipedia.org/wiki/Explicitly_parallel_instructio...
There's no reason for the VLIW machine not to get wider with time. The fact that Itanium didn't speaks more of its economic unhealthiness than of anything else.
Transmeta was 25 years ago and trying to apply vliw in a novel way to run x86 code, a daring idea that didn't pay off very well.
However, for that thankless and difficult task of creating an X86 clone (x87 was a nightmare) they received the great prize of competing against Intel in both manufacturing and marketing.
Moreover, they were a 32b machine in 2000 and 64b was right around the corner. Indeed when AMD64 came along, announced in 2000, it was simpler and forced Intel+HP to give up on Itanium. (Yes, Itanium had other problems but engineering problems can, to a degree, be surmounted. Market problems are harder.)
Also it's now mainstream (Intel does it) to have an intermediate code storage format with the introduction of trace caches. I wonder what the usual format of those is.
VLIW directly exposes hardware parallelism to the programmer, making more throughput available without requiring complex scheduling machinery (save die area + power). This works well for code targeting a single architecture (like embedded) where timing is predictable (like DSP code). Essentially VLIW machine code embeds a bunch of very specific assumptions about processor's internal design.
Here VLIW is alive and well for targets like a cell modem baseband or set top box video decoder chip that needs maximum efficiency and only runs the firmware from the manufacturer.
General purpose CPUs don't fit this model. They need to support running the same code across chip variants and generations, each with varying amounts of internal parallelism, cache, etc, and instruction timing is not predictable or deterministic (memory and IO latency, cache residency, etc etc). Modern high performance CPUs have extraordinarily complex and clever machinery to process what user space programmers think of as "machine code" to find and use as much parallelism as possible. This includes speculative execution based on behavior observed at run time, something not possible even in a perfect VLIW compiler. Given this you're not likely to see VLIW instruction sets exposed to user space. This is consistent with Linus's view I think.
Notably, as 3D accelerators became user-programmable VLIW stopped working well (see ATI/AMD GCN architecture rationale).
What Transmeta did is try to combine the advantages of VLIW hardware and stability of existing serial x86 code by filling the gap with a JIT living below the kernel ("Code Morphing Software"). There are examples like HP Dynamo that suggest this is a workable idea and in practice the 1st gen Crusoe product was decent. There are some different ideas why they failed but at least part of the problem is that they got squeezed by Intel's manufacturing advantage: 20 years ago being a year or two ahead in fab process was more effective than having a better design. I haven't seen any strong case that the technical approach wouldn't work at least as well today with TSMC and Samsung now leading in fab.
Interestingly the Denver architecture NVIDIA uses in their embedded products follows the Transmeta approach with an ARM front end and embedded JIT. Reportedly they were originally going to do x86 as well. I'm curious what's going to happen there now that they own Arm.
Last year's HotChips (or maybe even the year before) had a talk from a group that was working on a new design which is arguably VLIW (or has some elements of VLIW) and recently-ish they've been back in the news for having actual silicon in test. It's my hope that they are successful in their efforts, at least in some technical way.
Lastly somewhere in the back of my head is the vague idea that Apple doesn't actually use ARM cores either but instead uses a continuation the technology they bought with PA Semi only with an "ARM front end". I've long forgotten where I heard that from though.
He's also really practical guy if you happen to remember about Minix micro-kernel VS monolithic Linux-kernel ;)
Anyway, there are a couple of efforts which include VLIW tech which are in currently in the development pipeline now. So maybe in the next few years we will see some interesting things with it.
Snapdragon based phones have multiple Hexagon DSPs, they're VLIW that are also a commercial reality.
It's fair to say that there's a lot fewer folks who know how to write hexagon assembly than other architectures but you could play this optimization game w/it too.