The Convey HC-2 Computer – Architectural Overview [pdf]
conveycomputer.com
conveycomputer.com
It is not sufficient to recompile your C/C++ code to gain the performance advantages, you have to redesign the data structures in your code to match the characteristics of the hardware. For many code bases, redesigning all of your data structures is tantamount to a rewrite. This is not something a compiler can do for you automagically.
A follow-on caveat is that to design optimal data structures for these types of architectures, you need to be comfortable with understanding how silicon actually works and know how to do microarchitecture specific optimization in high-level code. Someone who has these skills can usually squeeze out several times the performance of typical software code on more vanilla CPUs, which closes the performance gap quite a bit and for some codes the latest greatest Intel CPU will actually be faster than a hybrid architecture.
The thing to keep in mind is that most benchmark comparisons I've seen for these types of architectures are (1) narrowly selected for workloads where they excel and (2) usually compare naively optimized CPU code with expertly optimized coprocessor code. In reality, due to skill availability you often see the reverse with expertly optimized CPU code and naively optimized coprocessor code that virtually erases the apparent performance advantages.
The major hurdle for exotic coprocessor architectures is that expert code designers that know how to exploit and use these architectures are incredibly rare. Consequently, you rarely see what they can actually do. Intel's Xeon Phi coprocessor has had similar issues.
Of course until actually using such tools, it's hard to tell. But the fact micron, a huge company which most likely is interested in big businesses ackuired convey, is a good sign to the use of such tools for the mass market and not just a bunch of highly trained experts.