Tachyum starts from scratch to etch a universal processor
nextplatform.com
nextplatform.com
Part of the problem with GPUs (and to a greater extent FPGAs) is that the toolchains are often terrible, buggy and opaque. They also make it really hard for people to write easier abstractions on top of them. I guess CUDA does do better along many of those axes than alternatives at the cost of being vendor-specific, but anything for this would be vendor-specific too.
So much code we write doesn't take advantage of GPU power because it's harder to program for and also you pay a latency cost for transmitting your data to/from the GPU. If this architecture makes GPU-style programming easier in that you just switch to using a different style of programming in the middle of your code and the CPU just uses different instructions without a big latency penalty that would be very cool.
As far as open sourcing the compiler, I suspect that having good documentation about what all the low-level instruction do would be as important if not more important than compiler source code. Just the compiler code wouldn't tell you why they do a given set of operation.
Moreover, if a group is designing the chip and compiler together, you may wind-up with a situation where they only know the compiler stuff works, they don't know what happens if you do various other things.
Be careful what you wish for.
Writing compilers (and hand-coded assembly) for theses types of 'poison bit' architectures is not for the faint of heart.
I've never heard that phrase before. Got any reading material on it?
Now the Moore's is basically at its end and x86 is partially stagnating, at least from Intel, other platforms like ARM are gaining traction and it seems like a good time to revisit VLIW.
I think another key factor is that most applications now run on top of platforms/frameworks rather than at the native level. This means you only need to port Linux, the JVM, node, python, and a few others and you captured a pretty large potential audience. Compare this to the mid 00's when moving to Itanium meant porting all you native apps.
It also failed because of Intel's very poor handling of the developers who wanted to switch to the Itanium architecture and eventually gave up because there was only support for big shops.
Likewise, I hope that this new CPU becomes easily available to enthusiasts and think that its viability even would depend on it.
The usual problem with these sorts of CPU microarchitectures for general-purpose computing is that they can't absorb variable cache/memory load latency. How is this one any different?
There is no ordinary compilation scheme that will solve this, even with complete omniscience, since the same function with different arguments will observe different latency. Maybe some magical feedback-driven JIT could do it, but that was tried in the Itanium era and never really worked either.
I can believe than a new VLIW processor can indeed perform well on HPC and ML workloads. But that doesn't sound particularly "universal" to me. Will people get good performance running relational databases on it? Graph algorithms? Compilers? Existing Java applications?
Itanium was never particularly impressive compared to its (roughly) contemporary competitors on the (roughly) same process, why would Tachyon be any better?
You say that tradeoffs change as a function of the process. E.g. wires becoming ever more expensive compared to transistors. Is that enough to tilt the playing field in favour of VLIW? I can see VLIW having an advantage here if the VLIW instruction bundle lines up 1:1 with the hardware pipelines (less routing inside the chip), and the workload is static so you can use profile guided optimization to work around the lack of dynamic scheduling (OoO).
However, once you try to create a VLIW ISA that maps to several generations and/or high/low end implementations you lose that 1:1 mapping, and for general purpose code compile time scheduling isn't particularly good. As Intel found out.