Personally, I found that the slide-ins and wobbly text made it difficult to keep my place. I tried Readability, but it didn't do well on it. Scrolling to the bottom to let the glitz run its course, and then scrolling back up to the top to read helped in the end.
Apologies for being critical, I wish you the best.
In comparison to other architectures, we have chose to stick to RISC, instead of some crazy VLIW or very long pipeline scheme. In doing this, we limit compiler complexity while still having very simple/efficient core design, and thus hopefully keeping every core's pipeline full and without hazards. The idea is that we just want to have a bunch of very simple and focused SPMD cores, so that we can have a MIMD/MPMD chip.
We are currently fixing the bugs on our single core FPGA demo, and hope to have our full 256 core cycle accurate simulator done by ~January/February. We want to release that (and our currently very early compilers) to the public ASAP.
>Local scratchpad memories are physically addressed as part of a flat global address space.
So from the programmers perspective each core will have a block of the address space, I.E.: 0-255, 256-511, 512-767, 768-1023 etc.?
Or is there address translation between units? Or if a thread is just built to arbitrarily execute on a unit, it'll have to pre-process its position for name space translation?
Also is there a memory locking in local scratchpad? (I maybe reaching).
A programmer will have full access to be able to handle memory however they want, but we want to be able to build out the tools to allow for a programmer to treat it similarly to an L1 cache (that is part of a shared memory space... that has different access times)
How does this compare to other multi-core processors / architectures e.g. epiphany http://www.adapteva.com/introduction/
have you ever heard of Epiphany[1] ? They claim to achieve 70 GFLOPS/WATT. Also the processor seems to be fairly fast and they manage to put 4096 cores on a chip. Though no activity in recent time. Maybe you could find some collaboration points with them..
[1] http://www.adapteva.com/epiphany-multicore-intellectual-prop...
The business case for open sourcing the ISA is that we want others to be making compatible chips. As a small startup with a new architecture, it would be GREAT if others were to make competing chips, as it would only further the architecture and software ecosystem, making it more competitive with the existing market incumbents.
To us, our floating point unit is our "secret sauce", but as an open source enthusiast, I don't want that to be locked up forever. My general idea right now is that we want to be able to tape out our first chip (at least the prototypes for it), and will open source the HDL for the non "secret sauce" parts of it. As we get onto further generations, I do want to open source the full design of our previous chips for free/open use.
The Epiphany 4's interfaces are 1.5GB/s serial (12.5Gb serdes), and there are 4 of those per chip, giving you an aggregate bandwidth of 6GB/s.
If you can spend a bit more power (or wait for 14/16nm process), we think it is possible to double that to 96GB/s per interface. Using some more exotic methods (which are purely in an idea stage right now, and not tested, are are a couple years away at best), we think it is possible to get that up to 128GB/s per interface.
Relevant paper (by Intel Research, actually)... there are quite a few differences (we are keeping it a lot simpler on the tx and rx ends, but that is some of our secret sauce... I can talk about it offline if you are really interested) http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=648778....
How much is the scratchpad memory in each of the processors?
If we can go to 14/16nm in the future, we are planning 512 and 1024 core versions with different amounts of memory depending on if you are memory or compute bound.