The NCube was another power of 2 machine, and could have up to 1024 CPUS, 64 per meter-square card. Each CPU was conventional, about 1 MIPS, with, I think, 128KB RAM. Stanford had one with one card of 64 CPUs. Well, 63; one was broken. It was a donation from an oil company that found it wasn't useful for seismic data processing. Nobody found a good use for it at Stanford, and it was donated to some other school.
Each CPU had a 0..1023 address. Message sending XORed the address of the source and destination, and the 0 bits told which paths would take the message closer to its destination. In theory there was supposed to be a path for each bit difference, like the Connection Machine, but actually there were some shared bus-type paths, I think. This allowed doing the whole thing with printed circuit cards and a printed circuit backplane, rather than a huge number of discrete wires. So the whole thing fit in a 1M cube.
With all these strange machines, the problem has to fit the machine, or you don't get the parallelism. Trying to fit existing problems into those forms was mostly a flop. Which is why general purpose shared memory multiprocessors with caches won out.
Massive parallelism seems to be problem-first. Graphics has a mostly-forward pipeline, so GPUs have a mostly-forward pipeline. Bitcoin miners have barely any intercommunication needs at all, so they're standalone things with minimal intercommunication. Back-propagation has a very standard form, and now we're seeing hardware that's sort of like a GPU but with lots of short multiply/add hardware.
What's striking is how much compute people have been able to get out of GPUs. They are sometimes the wrong tool for the job, but they're mass produced and affordable. The multi-million dollar one-off parallel machines, though, were mostly dead ends.