One of the most expensive parts of the hardware is memory, and fast memory is a lot more expensive to produce than slower memory. So we have the choice between using the same (and thus slow) memory throughout the system, or combining different kinds of memory so that at software has at least the chance
to run faster. This is a fundamental issue, and the only thing you can do is trying to find the optimal share for each kind of memory.Although all modern high-performance (edit: I mean non-embedded-microcontroller) computers work this way, it's not the only possible way. The Tera MTA takes a different, cacheless approach.
First, the problem with modern RAM in desktop machines is not that it sucks at bandwidth. You can get your bandwidth arbitrarily high by multibanking. Multibanking requires more buses or point-to-point links, but that's a tolerable cost.
The problem with modern RAM is that it sucks at latency, compared to what the CPU would like. Well, what do you do about latency? You make your requests earlier, and make sure you have other things to do in the meantime, when they get back. The Tera did this by having 128 sets of registers – 128 hardware threads — and switching to the next thread on every cycle. That means that, if all the thread slots were full, every thread only executed an instruction every 128 cycles, which is plenty of time to hide the latency of a slow memory fetch, as long as the memory bandwidth was adequate.
So basically every thread gets to pretend that it's running on a machine with zero-latency RAM — memory that's as fast as the registers. And pointer-chasing becomes as fast as looping over an array.
There are some other advantages to this design. Pipelining logic is very simple, because unless your pipeline gets insanely deep, you never have two instructions in the pipeline from the same thread, so you don't have register hazards.
(Cache is still beneficial in such a design, since it reduces the bandwidth that the links to main memory need to support. But the Tera didn't use it.)
I don't really understand why the Tera MTA failed in the market, and I suspect the problems were commercial rather than technical — customers had to take a big risk by porting their software to an unproven HPC platform, a platform whose performance characteristics were completely unlike anything else in the market (and unlike anything you can buy today). So customer uptake was insufficient to provide the cash flow needed to keep updating the design to keep up with Intel and AMD.
The technical reason such a design might fail would be if the silicon resources needed to support an entirely independent core were comparable to the silicon resources needed to support a hardware thread. Consider the GreenArrays GA144 chip: 144 independent cores, each with a tiny amount of independent RAM, on the same chip. Such a chip will be at least as fast as a chip with 144 independent register sets, but a single execution pipeline — in the worst case, it's bottlenecked on getting data out of RAM, and only one of its cores is usable, making it just as fast, while in the best case, it runs 144 times as fast. So a chip with 144 register sets needs to be cheaper — i.e. smaller — than the GA144. (Well, or easier to program, but presumably you can use more mainstream multicore chips to prove the example instead.)