Microchips That Shook the World
spectrum.ieee.org
spectrum.ieee.org
I think they were more influential to computer architecture than the original RISC and CISC ideas.
That's only 53 years in production now!
Wide network of peers with hands-on experience, easy access to tools (cheap devboards) can be a huge factor for hobbyists.
Yes, MS software was terrible, but this had little to do with x86 after the 32 bit transition.
This list is missing RF chips, though I suppose success is from RF technology on chips, not any single killer chip.
I'm hoping The Mill folks can bring back attention to VLIW.
The great wide hope lives on, somewhere in the future.
[1] https://people.eecs.berkeley.edu/~kubitron/cs252/handouts/pa...
I think the Mill people should concentrate on what VLIW has had some success with in the past: embedded. There will be tears if they go after general purpose.
The first Itanium sucked, and had a very bad memory subsystem. Itanium 2 was pretty competitive in performance, with one major exception: the x86 emulation. Why buy Itanium—even if it's very fast—if you could buy cheaper AMD64 which supports your existing software? Add in that Itanium compilers were few in number, all proprietary (and expensive), and didn't generate very good code for the first few years. Also, Itanium violated some programmer assumptions (very weak memory model and accessing an undefined value could cause an exception).
Now, Mill has a better way of extracting ILP than Itanium does, compiler technology is much better, and JIT compilers are very common. VLIW processors can be very good JIT targets. Mill, if it ever materializes, has enormous potential.
However the Mill doesn't extract ILP. Compilers do that. And given a code base there's only so much available. Yes, compiler technology is much better now and still there's only so much ILP available.
Lastly, VLIW has been tried with JITs at least twice: Transmeta Crusoe and IBM Latte. VLIW code generation is hard and it's harder if you have very little time which is the nature of JIT.
Ivan, if you're reading, think embedded.
[1] https://www.extremetech.com/computing/174023-tegra-k1-64-bit...
And the VLIW has to compete against that. There's only one place where VLIW has shown to have a survivable advantage and that's embedded.
https://www.hotchips.org/wp-content/uploads/hc_archives/hc26...
Also, the ee380 presentation: https://www.youtube.com/watch?v=oEuXA0_9feM
Yeah, Denver is in-order superscalar and it's a JIT but it isn't VLIW. Sad to say, they've tried JITing to in-order superscalar as well. They had a design win with the Nexus but even now NVidia is switching over to RISC-V for Falcon.
It is not from lack of effort that the JIT approach hasn't really worked. It's competitive but not outstanding. Denver, from the EE380 talk, thrives on bad bloated code. It's not so good on good code. This is not a winning combination.
Bad bloated code is 90%+ of the code in the world ;)
Also, shoot me an email at sixsamuraisoldier@gmail.com you seem to have some inside info on Denver, I'd love to chat (don't worry, I won't steal any secrets ;) )
[1] apparently it can sustain 5 instructions per clock given some very specific instruction mix.
[2] or out of order, really, Power8 being the exception.
Better compiler tech in the past few years (you mentioned JIT for example, which Denver has adapted to do quite well) has made VLIW a strong technical contender in several markets. Alas, the overhead cost of OoO is no longer the issue in modern computer architecture.
[1] https://www.extremetech.com/computing/174023-tegra-k1-64-bit...
[1] http://www.anandtech.com/show/8701/the-google-nexus-9-review...
The Denver 2.0 CPU is a seven-way superscalar processor supporting the ARM v8 instruction set and implements an improved dynamic code optimization algorithm and additional low-power retention states for better energy efficiency.
https://blogs.nvidia.com/blog/2016/08/22/parker-for-self-dri...Regarding the Mill: I still haven't heard a (good) answer for the perennial question: Where's the ILP?
Every time I ask this question, I'm pointed to the "phasing" talk. The issue is:
1) It's quite similar to a skewed pipeline (which Denver already has) in which you "statically schedule across branches" and delay the start until it's inputs are ready. Now, Denver did quite well indeed, but it's hardly a replacement for OoO.
2) Even from the examples that they show, it's clear that this is no where near 3x the ILP, at best, even from their own example, it provides 2x.
I hope the Mill can do well in embedded because they have quite a few good ideas (albeit many of which have already been done which they claim is their own, such as XOR-ing addresses, done by Samhita), and because the computer architecture industry is currently starved for innovation.
If so, it's on pseudo-vectorial instructions made by concatenating many opcodes on the same instruction word. I don't think they claim it performs better than OoO, just that it uses less power.
What? What do you mean by pseudo-vectorial instructions? Do you mean loop auto vectorization?
Just dismiss it. I was even going to delete it, but I took too long.
When a new technology or system is oversold and overhyped then underperforms (which is almost inevitable because of product maturity issues) people tend to respond sharply negatively.
1. There is absolutely no flexibility in scheduling. Any EU stalls due to memory access delays, non-constant-time instructions, or the like delays completion of the entire instruction word. This makes cache misses disastrous.
2. If a newer CPU makes improvements to the architecture (say, adding new EUs), programs cannot take advantage of those improvements until compilers are updated to support those improvements, and the programs are rebuilt with the new compiler. This is unlike typical superscalar architectures, where a new EU will be used automatically with no changes needed to the program.
1. It's not strictly true that there is "absolutely no flexibility in scheduling", there are a few techniques that can at least mitigate this issue, such as Itanium's advanced loads and a few others. On the aggregate though, the lack of MLP is absolutely what killed VLIW, especially nowadays.
2. This is true, but IBM's style (which the Mill has adopted) of "specialization" can effectively negate this issue.
Note the high end and general purpose, if you're talking about smartphone cores or supercomputing, then VLIW can actually perform quite well, and indeed Itanium enjoyed decent success in supercomputing and the Qualcomm Hexagon, albeit not really general purpose, finds itself in many processors doing solid work.
It's a tempting way to do things, it's very difficult and expensive to spin up and maintain dual development and production pipelines in parallel. But more often than not that's often the best way to go about things. Sometimes you make missteps that seemed sensible at the time. Sometimes you torpedo your whole business because you take too far a step back and you can't even survive long enough to let the new thing mature. Look at Intel, they've had multiple missteps that did tremendous damage to their business. They jumped on the netburst architecture bandwagon and that turned out to be a dead end. The Core architecture basically grew out of their Pentium Mobile work.
- National 32032 CPUs. Basically a VAX-on-a-chip, they were relatively slow and pretty buggy. (They were one of the processors we considered using in the Atari ST, but the 68000 won out, thank goodness).
- INMOS Transputer. One of the early "just apply +5V and ground and happy computing!" systems with zero support chips needed, quite easy to put into a grid of compute elements. Very interesting nibble-based instruction set, worth studying. Unfortunately the high-level language INMOS was pushing (Occam) was really strange, and Transputer performance was never really great. There were also some cool microcode bugs (you could lock a Transputer up for minutes by executing a bit-shift instruction with a very large count).
- RCA 1802. Another early microprocessor. Not really a failure, but not a huge success, except for its radiation-hardened variants, which have flown on many space missions.
- Just about any floating-point coprocessor chip. Ugh. Thank goodness the early Pentiums stopped that madness [1].
[1] I heard that Intel wanted to charge for doing floating-point operations on Pentiums. The idea was that you'd pay to have an "enable" fuse blown, which would purchase (say) 100M operations before a disable fuse got blown. There'd be something like 64 fuses, and the last one would enable floating ops forever. The rationale was that the only people who really cared about floating point were folks running spreadsheets . . . and those people obviously had money and would pay for performance! People hated this idea; then 3D gaming started to be popular (with Quake, et al) and suddenly flops were in general use and it would have been a really awful marketing move by Intel.
Encore did make one product which had fairly wide use: their Annex terminal servers- sold to Xylogics.
The idea of a separate co-processor that only did floating point ended with the 387. The 487SX (that was the coprocessor for the 486SX) was actually a 487DX (i.e. a full 486, with an FPU) that took over the system entirely and disabled the 486SX. But that whole scheme was still its own special kind of madness.