In-Depth View of Wave Computing’s DPU Architecture, Systems
nextplatform.com
nextplatform.com
Also putting all the "tensors" (I assume this means activations and weights?) into the external memory is downright dumb.
Only having 16MB of memory on die is also probably not a wise decision.
The use of GALS is certainly interesting though.
I do think the criticisms are valid regardless though. Data movement is far more expensive than computation, so having so little on-die memory is almost definitely a bad decision.
There seems to be an opening in the market for truly general purpose processors having say > 16 or 32 cores. So it would be nice to have a curated list somewhere of multipurpose chips.
The Adapteva Epiphany or the Rex Neo is probably closer to what you're looking for. However, once you ask for anything that handles control flow well, you pretty much end back up at the general purpose CPU, which is highly inefficient for a reason. For example, even your GPU today will become very inefficient with control flow due to branch divergence.
Quoted for irony of including MATLAB in a comment hoping for more open availability.
The execution policies that are provided with the stand will supprt launching normal OS thread on CPUs, but others could do all kinds of crazy things. They could perform the required operations on CPUs, GPU, Cuda, OpenCL or potentially on crazy hardware like this, where it makes sense to do so.
Here is one example, std::find has overloads that accept this: http://en.cppreference.com/w/cpp/algorithm/find
This is far from the only data-flow architecture out there (my favorite is the TRIPS instance of EDGE), but so far none have succeeded in the market.
The self-timed part is using NULL Convention Logic (NCL). Wave have dubbed their implementation of this WTL (see https://news.ycombinator.com/item?id=11469749). There are some very interesting questions about exactly how they use this, but it sounds like the synchronization between compute nodes is using NCL, but the compute nodes are conventional synchronous logic, clocked by the completion signal from NCL. That would certainly be an interesting hybrid, but the 6.7 GHz results begs a lot of questions.