Amazon's homegrown Graviton processor was very nearly an AMD Arm CPU
theregister.co.uk
theregister.co.uk
Granted, x86 has seen far more optimization work than ARM, but its unfortunate (but not unsurprising) they can't even come up with 1/3 of the speed of a 2016 Xeon. Amazon will need to rethink pricing, if they don't want this to get pigeonholed into running ARM-specific tasks only.
[0] https://twitter.com/david_schor/status/1067456537264832512
https://www.nextplatform.com/2017/11/08/qualcomms-amberwing-...
They have a pretty solid looking team, lots of former Intel folks [3]. I seem to recall reading they had several hundred employees in total, and at least several hundred million in funding, though I can't remember where I read that.
It would have been far more impressive for Amazon to launch with a product like this, or Cavium's ThunderX2.
[0] https://fuse.wikichip.org/news/776/x-gene-3-gets-a-second-ch...
[1] https://cdn.amperecomputing.com/documentation/hardware/eMAG/...
[2] https://www.theregister.co.uk/2018/09/18/ampere_shipping/
Ultimately it's the total cost to run your application that's important, not the raw speed of a particular core. So 1/3 speed is no problem if it's 1/4 the price (indeed this is the kind of trade-off you'd hope for in an ARM server).
Sadly from the same tweet thread that quote came from: 'All in all: Amazon having Arm-based offering is a great first step, but it does not seem to be a particularly appealing one at the moment, (at least for me: especially since the cost difference is nonexistent if I was to match my existing performance.)'. So, for this particular application at least, the A1 instances aren't giving a good cost/performance ratio.
Not necessarily. A core that's 1/3 the speed can multiply your response time by 3x. https://static.googleusercontent.com/media/research.google.c...
It's a bit like answering e-mail on a Celeron Chromebook compared to your dual Xeon Platinum quad Nvidia desktop. You won't notice the difference. 90% of the time the only difference I notice between my light-and-slow laptop I use for meetups and conferences and my heavy-and-beefy one is that the puny one has no built-in fan.
When running Python code, my very methodologically questionable benchmarks indicate the size of the L1 cache makes a huge difference, so I'll take an ARM64 box with a Core i9 worth of transistors used as cache over the i9 with an ARM64 worth of cache any time.
The only contender here is EPYC2.
Web developers are too busy developing their scripts and webpage. They won't have the time to drop down to the C intrinsic or Assembly level to speed things up.
Furthermore, the problem could be fundamental to this chip. There is only a 1MB-per-core last-level cache (L2 cache). That likely isn't enough for anyone running your Linux / Database / PHP / Web Server stack. The chip does outstanding on "small" benchmarks like C-Ray, but it does horrible in practice, where you have more code (probably more than 1MB of "hot" code) running on the chip.
https://twitter.com/eastdakota/status/976560820611031040
Presumably ~1/3 less power means AWS gets a better margin on these.
Yip, at 4x the core count it is only 2x as fast. I also suspect that 2x RPi's would be less power usage than one of these platforms.
But then this is it's first generation dip into this area and they will be seeing what people actually want to run and will fine-tune their custom silicon from that for version 2.
As for pricing, they need to make this more tempting so they can get the metrics they need to see what version 2 should focus upon.
What I'm wondering is would such a chip marry up with an FPGA and present a more flexible and appealing option.
[EDIT ADD] It would appear this set of tests are heavy on single threaded usage and with that, fair results and like many, got taken in on the jonalistic spin. So it's in this test effectively comparing a single core, so mentioning it is a 16 core cpu compared to a RPi, does twist how the results are viewed.
That's almost certainly wrong information. Cortex A72 is an out-of-order CPU unlike Cortex A53 in RPI3. Also Graviton is running at 2.3 GHz vs 1.4 GHz for RPI 3b+.
You'd expect 16 core Graviton to outperform RPI3b's 4 core A53 by about 8-15x.
I have looked at some other benchmarks, and it looks like this offering only pays off on certain tests (N-Queens being one), how that pans out with real-world usage is what we will find out over time. Though the aspect that even in a few tests that this stands out, makes it viable for some real-world usage. Time will tell and at least even if it pays of in 1 out of 10 tests, those tests will translate into real usage. But one shoes does not fit everybody, but always good to have a choice and for that, this is good.
I don't think the goal of ARM server chips should be core-for-core near-parity with Intel big cores anytime soon, because those have much bigger die area and power needs, so you potentially pay in number of chips per wafer, yields, density, and power use. Instead you keep something reasonably small, try to get single-thread perf good _enough_ for a chunk of apps, and try to provide much better cost per unit throughput.
(They could also go for efficiency improvements to further cut power-related costs or make up for more consumption by faster cores. These seem far-out but theoretically: mobile-like big/little arrangements so light loads use even less power, or boxes with low-power RAM, or (least far-out) more of their Nitro IP moved onto the SoC. Who knows, just saying it's a wide-open design space.)
If this takes off enough for Intel to really pay close attention (not a certainty!), they do have their own cool scale-out experiments in Xeon D (an SoC Facebook originally wanted with relatively few cores and low clocks, with some SKUs including compression/crypto acceleration) and their Atom server chips. Getting something good and Intel responding with something better isn't the worst possible result.
I'm more surprised it isn't roughly 1/2 the price it's being pushed at. It seems to be more expensive than the x86 side.
https://www.phoronix.com/scan.php?page=article&item=ec2-grav...
It warrants its own discussion.