Startup to Open-Source Parallel CPU
eetimes.com
eetimes.com
I've worked in a start-up company similar to your company.
We developed a 256-cores RISC processor, only shared memory was used between all the cores instead of a mash-up of a memory block for each core and DMA for transactions.
How do you intend to synchronize work between the different cores? How a compiler will abstract away the memory synchronizations? Which programming language is going to be used? So many questions as this is such a complex area in computing...
From my personal experience of over 5 years developing such chip in a start-up company, the cost of production will probably be a huge obstacle. Good luck!
Technically, you don't need to synchronize the cores... it's a MIMD/MPMD system that is not in lockstep. It is up to the programmer (with help from the compiler) to make sure you don't do anything too stupid ;) As for the programming languages, C and Fortran are the big ones to us. We hope once our LLVM backend is improved, you'll be able to run anything you like on the cores themselves. As for the programming model, the three we like the best are CSP (Go), Actor (Erlang), and PGAS (C, Fortran, Chapel, a few other research ones). If you're familiar with SHMEM, that's the closest thing we can think of currently.
When it comes to cost, we're trying to develop as much ourselves to reduce licensing costs. We've been able to do a pretty good job (if I do say so myself) as just two people so far with no capital. Fabrication costs are the killer, with it being about ~$500k per shuttle run, and $5m-$7m for a mask set when we actually go to full production.
All our cores shared the same memory for both data and instruction and we based our synchronization of work-load on a hardware instead of software. It yielded such a huge speedup for execution time that most of the companies simply ignored our results as fake. :)
Choosing a programming language is crucial - we went with C and a declaration language for tasks. Today (4 years after we closed the company) I am not sure whether it was the best decision. The simpler the parallel definition in code the better. Programmers as getting confused easily.
Apologies for being critical, I wish you the best.
In comparison to other architectures, we have chose to stick to RISC, instead of some crazy VLIW or very long pipeline scheme. In doing this, we limit compiler complexity while still having very simple/efficient core design, and thus hopefully keeping every core's pipeline full and without hazards. The idea is that we just want to have a bunch of very simple and focused SPMD cores, so that we can have a MIMD/MPMD chip.
We are currently fixing the bugs on our single core FPGA demo, and hope to have our full 256 core cycle accurate simulator done by ~January/February. We want to release that (and our currently very early compilers) to the public ASAP.
How much is the scratchpad memory in each of the processors?
If we can go to 14/16nm in the future, we are planning 512 and 1024 core versions with different amounts of memory depending on if you are memory or compute bound.
>Local scratchpad memories are physically addressed as part of a flat global address space.
So from the programmers perspective each core will have a block of the address space, I.E.: 0-255, 256-511, 512-767, 768-1023 etc.?
Or is there address translation between units? Or if a thread is just built to arbitrarily execute on a unit, it'll have to pre-process its position for name space translation?
Also is there a memory locking in local scratchpad? (I maybe reaching).
A programmer will have full access to be able to handle memory however they want, but we want to be able to build out the tools to allow for a programmer to treat it similarly to an L1 cache (that is part of a shared memory space... that has different access times)
The business case for open sourcing the ISA is that we want others to be making compatible chips. As a small startup with a new architecture, it would be GREAT if others were to make competing chips, as it would only further the architecture and software ecosystem, making it more competitive with the existing market incumbents.
To us, our floating point unit is our "secret sauce", but as an open source enthusiast, I don't want that to be locked up forever. My general idea right now is that we want to be able to tape out our first chip (at least the prototypes for it), and will open source the HDL for the non "secret sauce" parts of it. As we get onto further generations, I do want to open source the full design of our previous chips for free/open use.
The Epiphany 4's interfaces are 1.5GB/s serial (12.5Gb serdes), and there are 4 of those per chip, giving you an aggregate bandwidth of 6GB/s.
If you can spend a bit more power (or wait for 14/16nm process), we think it is possible to double that to 96GB/s per interface. Using some more exotic methods (which are purely in an idea stage right now, and not tested, are are a couple years away at best), we think it is possible to get that up to 128GB/s per interface.
Relevant paper (by Intel Research, actually)... there are quite a few differences (we are keeping it a lot simpler on the tx and rx ends, but that is some of our secret sauce... I can talk about it offline if you are really interested) http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=648778....
Personally, I found that the slide-ins and wobbly text made it difficult to keep my place. I tried Readability, but it didn't do well on it. Scrolling to the bottom to let the glitz run its course, and then scrolling back up to the top to read helped in the end.
How does this compare to other multi-core processors / architectures e.g. epiphany http://www.adapteva.com/introduction/
have you ever heard of Epiphany[1] ? They claim to achieve 70 GFLOPS/WATT. Also the processor seems to be fairly fast and they manage to put 4096 cores on a chip. Though no activity in recent time. Maybe you could find some collaboration points with them..
[1] http://www.adapteva.com/epiphany-multicore-intellectual-prop...
Other than the grid interconnect, how does the architecture differ from Xeon Phi? In particular, what allows it to get such dramatically higher power efficiency? I'd have guessed that Intel's smaller process would make it difficult to match.
[1] It feels awkward that this sentence is included twice on such a short page.
As for comparison to the Phi... it is the fact that our actual core size (and thus the whole chip) is MUCH smaller. The Phi's cores are actually based on the original Pentium architecture (they are just downsized P54Cs) with added AVX instructions. Contrary to popular belief, they are not full x86-64 cores.
Intel has not released the official die size for the Phi, but has said it is about ~5 billion transistors (at a 22nm process), and independent "guesstimates" have pegged the die at 600-700mm2 (there is one place that says 350mm2, but that is false). For the top of the line 61 core Xeon Phi, it uses 300W, with a theoretical peak performance of 1.2TFLOP of double precision performance. That gets you to about 4GFLOP/Watt.
In comparison, our entire will be under 100 million gates, with each core (excluding memory) being around 100k gates. At a 28nm process, our core size (without memory) is a little under 0.1mm2. Our theoretical peak performance per compute chip is 256 GFLOPs double precision, while it should be using around 3 Watts, giving us a performance per watt ratio of ~85GFLOPs/Watt.
Intel has even said that their next generation Xeon Phi, made at their 14nm process, will be at 14 to 16GFLOPs/Watt. At SC14 last week, they made a soft announcement for the following generation at 10nm process will be in the ~2018 timeframe, and that is estimated at only being around 25GFLOPs/Watt.
The bottom line is Intel is just following Moore's law, and is sticking to big and complex systems, which while retaining legacy compatibility, kill you when it comes to efficiency.
Then again, I don't think most popular "RISC" CPUs are very RISC like (Does that make me a RISC hipster?). While making things general purpose, there is no reason for ARM processors to be OoO, do branch prediction/speculative execution, or any more of these crazy things... it reduces efficiency in the long run.
Take a look at this for instance: http://chip-architect.com/news/2013_core_sizes_768.jpg ... an AMD CPU and ARM chip are virtually the same size at the same process node. I find it insane that one of our cores is a bit less than 1/5th of the size of a cortex A7, and can do more FLOPs than it. Then again, we have focused on doing that, but still.
The fact that we are using mostly open source tools for our development (such as Chisel: chisel.eecs.berkeley.edu), our development time has decreased and productivity has had a huge boost compared to if we were just writing in Verilog. We also made the decision to not use off the shelf IP, which while difficult to verify, we think we make a much better system. Compare this to the Epiphany implementations, which while having a custom ISA and basic core components, used off the shelf ARM interconnect, off the shelf ARM memory compiler, and many other things that were used to minimize development time and verification. While this is a bit more difficult upfront, since we are keeping our components very simple, our verification is no where near as complex as a "normal" processor. Plus we don't have to pay $500k-$1m+ in upfront licensing fees.
How much scratch memory is there? Is it SRAM?
Are they licensing anyone's IP for the interconnect, or CPU? What's the bandwidth of the interconnect? Is is packet-oriented? How does fair-routing work?
We currently have 128KByte dual-ported SRAM per core (which is physically part of the core, and not a giant array somewhere else on die). It has single cycle latency to the core's registers and to the Network on Chip router.
The on chip mesh network is custom 128 bit wide going core to core. The router can do a read or write to SRAM per cycle AND allows a passthrough to another core in the same cycle.
Our chip-to-chip interconnect is a custom 64 bit (72 lane) parallel interface allowing 48GB/s. There are two of these (unidirectional) interfaces per side, giving you a total of 8 of these interfaces per chip.
http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=648778...
I myself have not taped out something on a modern process, but have advisors who have. My co founder and I do have nanofab experience, so we do understand the physical complexities of fabrication first hand.
The biggest problem (even with our solutions for skew and crosstalk) is just the number of pins/traces on the board, but that's not unsolvable... nothing a ~10 layer PCB can't solve.
How big chunks of computation you need to do in a node for this to be effective?
Our single precision numbers are actually double that of double precision (compared to a GPU or most other SIMD systems, which have independent FP32 and FP64 FPUs, we have a single combined unit). While our ISA is pure 64 bit, we have packed 32 bit FPU instructions for doing two single precision FLOPs per cycle.
Edit: didn't notice same was asked before too. Regardless, how many FMAs can be executed in the scenario I gave, also sending 16 bytes and receiving 16 bytes for each FMA?
Synchronization should be handled in a dataflow-driven program design. Any sort of mutex/semaphore/etc. will have to be software defined or interrupt-driven.
In regards to neighboring node being memory mapped... are you asking about another compute chip or another node (with the 16 compute chips + GaMMU)? All of the compute chips in a grid are part of the same flat global memory map, and have DMA capabilities between each other. Once you get out of the compute grid (that is managed by the GaMMU), then that is a separate memory space, but can be accessed through some other layer through something like MPI, for instance.
Can you do reasonably efficient [arbitrary size] fixed point arithmetic on your hardware? Do you have 64-bit add with carry and 64-bit multiply with 128 bit results? I'm interested in 64.64 and 128.128. Needed operations are add, sub and multiply.
I think compute grids like these could very well be an important part of computing in the future, maybe even the most important part. Ever since when I first saw an article about transputer. Grids or VLIWs, sadly software is always the pain point. I wish you luck, please get this working and right.