I think you're underestimating the requirements and mastery of cloud companies. Something like an Amazon lambda could virtualize 4 cores per instance and host 256 lambda execution units on a single chip. The use cases are endless
Unless the architecture has changed drastically from the earlier Epiphany, they can't be virtualised like that, and each core are way too slow to be suitable for lambda except for software written specifically to take advantage of the parallelism of the architecture.
You still need to recompile code for the new architecture, and taking full advantage of it wisely is not easy... but may be worth it in many use cases. Part of the problem is that it's not 100% clear which use cases these are and how to market it. Probably unit calculation per watt is the most likely performance advantage, but it's still amazingly hard to sell people on that sometimes
Some parallel algorithms will scale to bigger (more parallel) chips the way binary programs got more performance with clock higher frequencies. That's the holy grail..