The chip is about the same size as the Apple A10, so in terms of silicon area it's in the consumer domain, but price will only come down to consumer levels if shipments get into millions of units. Big companies take a leap of faith and build a product hoping that the market will get there. Small companies get one shot at that. With University volumes and shuttles, we are talking 100x costs. So the $300 GPU PICe type boards become $10K-$30K with NRE and small scale productio folded in.
It would be a BIG mistake to assume 1024 working cores. If you want to scale your software you should take a look Google/Erlang and others. Not reasonable to demand perfection at 16nm and below...
Not saying we won't have chips with all cores working, just saying you shouldn't count on it.
Sure, there will be evaluation boards, they just won't be generally available at digikey and won't cost $99. More information about custom ISA will be disclosed once we have silicon back.
The GDS is completely tied up in NDAs due to the foundry. The EDA combines/translates open source code with proprietary blobs to produce a "super secret" GDS binary blob that gets sent to the foundry for manufacturing.
Thanks for articulating. As you know, there is no right answer as it depends on workload. Now if we could only build a specific chip for every application domain....
Well, the Parallella has shipped to over 10,000 people and it still selling at Amazon an DK, so no the dream is not dashed in any way. The number of publications and frameworks around Parallella is growing every month...
No reason to drive a 1024 core chip to the broad market when most applications aren't ready to use 16 cores. With this chip we focus on customers and aprtners who have proven that they have mastered the 16-core platform.
Yes, you can call it scratchpad or sram. The point is that there is no hardware caching. The local SRAM is split into 4 separate banks so it is "effectively" 4 ported. DRAM controllers is up to the system designer. This is handled by the FPGA. (like previous epiphany chips).
I'll pass for now...Gaisler is in the business of consulting, we survive by building products. I am happy to release sources, but it's completely up to the EDA company.
[edit: was thinking of the wrong Gaisler, still will pass]
Agree, but people have all kinds of pre-conceived notions about co-processors so let's clarify some things: e5 can't self-boot, doesn't have virtual memory management, and doesn't have hardware caching, but otherwise they are "real" cores. Each RISC core can run a lightweight runtime/scheduler/OS and be a host.
I have two excuses for why RISC-V didn't make it it. My February RISC-V post stated that we will use RISC-V in our next chip. We were already under contract for this chip so I was referring to the next chip from now. I had hopes of sneaking it into this chip, but ran out of time. Both lame excuses, I know. I am firmly committed to RISC-V in some form in the future. For clarity, I am not talking about replacing the Epiphany ISA with a RISC-V ISA.
Paper stated that 500MHz number was arbitrary (had to fill in something for people to compare to). Agree that 500MHz with 16nm FinFet is ridiculously slow. We are not disclosing actual performance numbers until silicon returns in 4-5 months. 28nm Epiphany-IV silicon ran at 800MHZ.
Yes, we are leaving 2X on the table in terms of peak frequency compared to well staffed chipzilla teams. Not ideal, but we have a big enough of a lead in terms of architecture that it kind of works.
Hours were over a 12 month period, but yes...the pace was relentless. All ambitious projects, including many kickstarer projects get done because creators end up working for free for essentially thousands of hours. In this case, we were on a fixed cost budget so those hours were "my problem".
Not going to happen in the near term. There is no way to meet the price point needed to compete in the low cost SBC market with the Epiphany-V. Believe it or not, the $99 Parallella was priced too high to reach mass adoption.
It comes back to the programming model. Synchronization is all explicit. See publication list. Includes work on MPI, BSP, OpenMP, OpenCL, and OpenSHMEM. The work from US army research labs on OpenSHMEM is especially promising. It's a PGAS model.
That's an older paper, but yes there have been more than one independent study showing 25x boost in terms of energy efficiency. See Ericsson FFT paper, OpenWall bcrypt paper, and others at parallella.org/publications.
In general it was built for math and signal processing (broad field). Within those fields, more specifically it was designed initially for real time signal processing (image analysis, communication, decryption). Turns out that makes it a pretty good fit for other things as well (like neural nets..). Here is the publication list showing some of the apps. (for later, server is flooded now): http://parallella.org/publications
Agree with you. Taping out chips is still like going to the dentist for a root canal. Still, we may have to disagree regarding the absolute numbers. For every one of our chips so far, we made RTL changes less than 24 hours before tapeout. (clearly this means our basic blocks are very small and we don't have a lot of them).
Until 2011, all of our tapeouts were shuttles. (for our initial market the per die cost was fine). You get 50 dies/wafer on a shuttle run and can order a lot of wafers...with a theoretical per wafer price of $5000 that would mean $100 die, which is not bad for some systems.
I provided ranges in the paper. To focus in further, you will need to get actual quotes from vendors. This is because the price is completely opaque and depends on circumstances.
imho, the only reason to go to 14nm and below is if you are pushing the envelope in terms of performance (HPC) or integration (smartphone). Everything else can probably stay at 28nm or above and be combined with 14-7nm using 2.5D or 3D package integration. Note that unless you are a Tier-1 system player you will "never" be able to get known good dies from a big fabless semi vendor, but taping out twice yourself and combining them yourself is pretty straightforward if you have a good packaging partner (like Amkor/ASE)