The reasons why this almost never works is one of the following:
- They assume they can move hardware complexity (scheduling etc, access patterns into software). The magic compiler/runtime never arrives.
- They assume their hard-to-program but faster architecture will get figured out by devs. It won't.
- They assume a certain workload. The workload changes, and their arch is no longer optimal or possibly even workable.
- But most importantly, they don't understand the fundamental bottlenecks, which is usually memory bandwidth. Even if you increase the paper specs, like FLOPS total, FLOPS/W etc. youre usually limited by how much you can read from memory. Which is exactly as much as their competitors. The way you can overcome this is by cleverness and complexity (cache lines, smarter algorithms, acceleration structures etc), but all these require a complex computer to run with all those coherent cache hierarchies, branching and synchronization logic etc. Which is why folks like NVIDIA keep going on despite facing this constant barrage of would-be disruptors.
In fact this continue to be more and more true - memory bandwidth relies on transcievers on the chip edge, and if the size of the chips doesn't increase, bandwidth doesn't increase automatically on newer process nodes. Latency doesn't improve at all. But you get more transistors to play with, which you can use to run your workload more cleverly.
In fact I don't rule out the possibility of CPU based massively parallel compute making a comeback.