Chinese Chipmaker Unveils 64-Core ARM Processor
top500.org
top500.org
Hmm..
Process:Manufacturing with 28nm process
And there you go.
Not sure if people are aware, but China doesn't have any sub-28nm (or 22nm?) fabs. TSMC has some of course, but the Taiwanese government is very, very strong on keeping those plants in Taiwan (although they are fine with TSMC etc building less advanced fabs in China).
Until they get a process shrink on this, I think.. well, I'd like to see some independent benchmarks.
Edit: The Samsung Exynos 5433 (4 cores on a 20nm process[1]) maxes out at 3.78 GFLOPS[2] on selected benchmarks. Call it 1 GLOP/core - I find it unlikely that this thing is going to get 500 GFLOPS with 64 cores, even with a higher power consumption.
[1] http://www.anandtech.com/show/8718/the-samsung-galaxy-note-4...
[2] http://www.anandtech.com/show/8718/the-samsung-galaxy-note-4...
That isn't true, SMIC has a 28nm process(very low yields, and only fabbing designs from Qualcomm afaik, but supposedly does exist, as far as anyone from the west can reasonably ascertain). [1]
1. http://www.smics.com/eng/press/press_releases_details.php?id...
>That isn't true, SMIC has a 28nm process
I don't understand your comment. 28 is not sub-28.
These benchmark numbers you found about the Exynos 5433 are so low probably because "Geekbench 3" is not very good at reaching the maximum theoretical performance, as it runs a bunch of real-world computing tasks. In general you need to hand-code assembly loops of fused-multiply-add instructions to reach the max theoretical perf and clearly this is not what Geekbench does, look at the list of FP workloads it runs: http://support.primatelabs.com/kb/geekbench/geekbench-3-benc...
Edit: Though last year's announcement was just the planned specs, perhaps the unveiling here is the prototype itself and not the concept.
[Chinese links] http://www.ltaaa.com/bbs/forum.php?mod=viewthread&tid=364899 http://bbs.kafan.cn/thread-1849400-1-1.html
"FCBGA package with 2892 pins"
Yikes! Thats pretty dense on the other side, would love to see a picture.It's a 100 watts, so at 3.3 volts its drawing 30+ amps, so I'm guessing that a good chunk of those pins are power and ground. Is that a good theory?
standard Vcc for 28nm is 850-1050mV (varies based on the exact process)
so yeah more like 100A and probably that's all going through the pins without on die power regulation
Consider that the VRM's have as much or more silicon in them than in the host processor, and they have completely different breakdown voltage and switching speed requirements relative to a CPU.
Lets talk board layers. My best board was two data, one power and one ground. I was nowhere near 100 amps for current needs. How many layers in a motherboard to support this CPU?
I'm thinking that these state of the art chips are really pushing the support infrastructure of board layers and power supplies.
Thanks1
available in two forms, as a coprocessor or a host processorHowever perhaps this processor beats the Xeon Phi 7210 in perf/price as the latter is horrendously priced at $2438 (list price) as well as perf/watt.
And in large enough volume, assuming 50% gross margins, sure it could be cheaper.
Given that, my guess would be that Erlang/Elixir/LFE/$INSERT_OTHER_BEAM_BASED_LANGUAGE_HERE would do pretty darn well with a high number of cores. I hope one of these days I'll be able to afford such a machine (whether with this particular ARM processor or something else with a ridiculous number of cores/threads, like a modern POWER or SPARC CPU) or have access to one so that I can experience for myself exactly how darn well :)
It's also worth noting that Erlang has had a lot of design around clusters of independent nodes, which means even more extreme problems when it comes to data copying between nodes. I reckon intercore message copying is significantly more performant than internode message copying.
In contrast, BEAM's SMP support is automatic AFAICT; the scheduler will happily distribute processes across as many cores as it can access (1 BEAM thread per hardware thread by default, so a quad-core CPU with one thread per core and a dual-core CPU with two threads per core will both be loaded with four BEAM threads unless BEAM is configured to do something else).
Ponylang looks pretty cool, though; I'll have to check it out.
which compile Erlang to Native code.
http://wccftech.com/us-government-bans-intel-nvidia-amd-chip...
http://www.extremetech.com/computing/227059-amd-announces-ne...
Slides 12-16 http://insidehpc.com/2016/08/phytium-china-unveils-64-core-a...
Each 8-core "panel" has its own L2 and DCU The L3 is globally shared, with the usual cache coherence protocol. 30ns latency for L3 hit.
I am not an expert on cache performance, but this certainly seems like an up-to-date design.
Can anyone compare this to Xeon and Phi ?!
A price would also be nice.
It would be very interesting to see real-world benchmarks compiled independently.
On the other hand this is good for the future of computing in general. The worst case scenario would be for things to stagnate with no competition, offering no incentive for anyone to push beyond traditional Moore's law type scaling.