I have two, and they're fun little toys and I wish they'd have gotten further, with the caveat that actually making use of them in a way that'd be better than just resorting to a GPU is hard.
The 16-core Epiphany[2] in the Parallella is too small, but they were hoping for 1k or 4k core version.
I'm saying it's a similar idea because the Epiphany cores had 32KB on-die RAM per core, and a predictable "single cycle per core traversed" (I think, not checked the docs in years) cost of accessing RAM on the other cores, arranged in a grid. Each core is very simple, though still nowhere near the simplicity of a 6502 (they're 32 bit, w/with a much more "complete" modern instruction set)
The challenge is finding a problem that 1) decomposes well enough to benefit from many cores despite low RAM per core (if you need to keep hitting main system memory or neighbouring core RAM, you lose performance fast), 2) does ot decompose well into SIMD style processing where a modern GPU would make mincemeat of it, and 3) is worth the hassle of figuring out how to fit your problem into the memory constraints.
I don't know if this is an idea that will ever work - I suspect the problem set where they could potentially compete is too small, but I really wish we'd see more weird hardware attempt like this anyway.