Arm Announces Neoverse V1, N2 Platforms and CPUs, CMN-700 Mesh
anandtech.com
anandtech.com
> 49% of AWS EC2 instance additions in 2020 are based on Graviton2
Surprised at this level of Graviton2 adoption in AWS at this stage. Any clues as to who is using these instances?
Edit: Presumably Intel's shrinking Q1 2021 Data Center revenues are partly as a result of this.
Percentage of compute power would be cool to know here.
I host a few very low traffic sites & I'm in the process of switching from a basic DO Droplet to a pair of low-end Gravitons. Will save me money and give better peak performance for my workloads.
More likely some very big customers (peer comment mentions Twitter) moving to Graviton2 for cost savings.
Edit: Nope, not yet, but close..you still have to change the radio button: https://imgur.com/a/W0Sweyy
I'm having trouble figuring this out - a t4g.micro is $6/month, before any storage or data transfer costs. The roughly equivalent DO offering is $5/month, inclusive of 25GB SSD and 1TB transfer. Even with a reserve instance discount and significantly less than 1TB outbound transfer, DO seems likely to be cheaper.
It was both AMD and ARM.
There are many work loads that G2 offer immediate cost / performance advantage. AWS charges per vCPU, which is one thread on Intel/AMD and one Core on ARM. So you get ~30% performance improvement along with a ~30% lower cost for using ARM Graviton Series. Most of them have reported a total of 50% reduction in cost. For those that have hundreds if not thousands of EC2 running which fits that workload advantage, this is too much saving to pass on.
There are many SaaS running on EC2 that has mentioned their success on twitter and various other places.
Worth pointing out, this is with Amazon installing as many as they get from TSMC.
A few months ago on HN I wrote [1] about how half of the Intel DC market will be gone in a few years time.
Edit: Another point worth mentioning, this is as much of a threat to Medium and Smaller Size Cloud like Linode and DO where they dont have access to ARM (Yet). And even when they get it Amazon have the cost advantage of building their own instead of buying from a company ( Ampere ).
The question is what options do they have to deal with this?
Even a decade ago that would've been unthinkable, but today, making a cookiecutter SoC is relatively easy because nearly everything can be taken off the shelf.
Production costs though.... sub-10nm mask set costs completely rule out anything resembling a startup competing in this area.
I think 65nm was the last golden opportunity to jump on the departing train. It was still posible to ship a cookie cutter chip under $1m, now... no way.
Now, Semi industry is basically Airbus vs. Boeing.
Nuvia basically never intended to really compete Intel, or AMD heads on. Their $30m stash would've been just enough for a single "leap of faith" tapeout on a generation old node, and a year of life support after.
They were aiming for a quick sell from the start too.
I definitely don't agree with premise that it's now Boeing vs Airbus now (certainly less so than it was a few years ago when x86 was the only game in town).
Back in 65nm, 40nm days, big tapeouts were already costing in high 6 figure digits in masks.
And... masks are not the most expensive items on the signoff costs these days.
Specialist verification, outsourced synthesis, layout, analog, physical, test, and other specialist services will easily cost more than the maskset for <40nm.
I would not be surprised if tier 1 fabless already spend $10m+ per design just on them.
IoT/Edge deployments are less standardised than other computing workloads. Developers and integrators in this area already expect to deal with a lot of bother when working with a new chips. Also, the margin on these devices is usually razor thin, so the potential savings from not paying ARM licensing fees would be more appreciated.
Finally, RISC-V's modular approach allows for a greater level of flexibility and innovation, which will allow manufacturers to further differentiate and gain a competitive advantage. This is especially relevant for IoT/Edge solutions where thermal and power budgets are heavily constrained.
For EC2 we run on spot and spot c5.metal are cheaper per vcpu than c6g.metal, so we haven't prio'd benchmarking our compute loads.
I have evaluated going ARM, but I ended up deciding the savings were not worth it.
Not only you need to mantain 2 archs simultaneously for some time, but porting some stuff to ARM, (e.g. Python) can be a pain in the ass.
Finally, my devs work in AMD64 and that would be another source for "why does this work in dev but not prod".
I can see a use case for building a CI/CD pipeline on Raspberry Pi's.
AArch64 & arm
A number of new CPUs are supported through arguments to the -mcpu and -mtune options in both the arm and aarch64 backends (GCC identifiers in parentheses):
Arm Cortex-A78 (cortex-a78).
Arm Cortex-A78AE (cortex-a78ae).
Arm Cortex-A78C (cortex-a78c).
Arm Cortex-X1 (cortex-x1).
Arm Neoverse V1 (neoverse-v1).
Arm Neoverse N2 (neoverse-n2).
Good to see work going into this at the proper times. (Not that that was much of a problem for CPU cores in recent times. Still not a matter of course though.)https://blog.dbi-services.com/aws-postgresql-on-graviton2-aa...
https://github.com/microsoft/STL/issues/488
We are about 9-14 months away from the right pieces making their way through the software ecosystems where this will be almost a non-issue.
Exciting times for everyone!
Really the important piece for making distribution binaries not suck is ifuncs/multiversioning. But library and app authors currently are required to deliberately use them. Which is fine for manual optimizations that use intrinsics or assembly (and e.g. standard library atomics) but I'm not sure any compiler currently would automatically just do that for autovectorization.
- when you have a particularly CPU-intensive application, you'd hopefully compile it to target your system
- the cloud providers can just do a custom Debian/Ubuntu/... build for their zillions of identical systems
- the library loading mechanism on Linux is slowly getting support for having multiple compile variants of a library packaged into different subdirectories of /lib (e.g. "/usr/lib64/tls/haswell/x86_64")
Also I was mostly trying to point out as a positive how well the interaction is working there between ARM and the GCC project. I wish it were like this for other types of silicon.
(CPU vendors all seem to be getting this right, and GPUs are slowly getting there, but much other silicon is horrible… e.g. wifi chips)
DDR5,PCIE 5.0, SVE speedup and 40% IPC improvement put a big smile on my face.
Hopefully ARM on cloud will result in cheaper prices.
The great majority of developers use Windows or Linux according to every Stack Overflow survey from the past ten years. Only ~25% use a Mac.
If 25% of servers switch to ARM that is massive.
""" And the only way that changes is if you end up saying "look, you can deploy more cheaply on an ARM box, and here's the development box you can do your work on". """
(emphasis in original)
https://www.realworldtech.com/forum/?threadid=183440&curpost...
Thanks, I did not know about this!
However I'd be amazed if they don't release some kind of managed service for running Swift code in the cloud. Caveat emptor, though.
I guess, he didn't expect this to take well after his retirement.
Also, the complex memory operands can be executed directly because you can add more ALUs inside the load/store unit. ARM also has more types of memory operands than a traditional RISC (which was just whatever MIPS did.)
Why do you think they have good conequences?
The upside to variable length instructions is that they are on average shorter so you can fit more into your limited cache and you make better use of your RAM bandwidth.
The downside is that your decoder gets way more complex. By having a simpler decoder Apple instead has more of them (8 wide decode) and a big reorder buffer to keep them filled.
Supposedly Apple solved the downside by simply throwing lots of cache at the problem and putting the RAM on-chip.
I'm not a CPU guy and this is what I've gathered from various discussions so I'm happy to be corrected.
There he analyses existing RISC and CISC architectures, and counts various features of their instruction sets. They clearly fall into distinct camps.
But!
Back then (mid 1990s) x86 was the least CISCy CISC, and ARM was the least RISCy RISC.
However, Mashey's article was looking at arm32 which is relatively weird; arm64 is more like a conventional RISC.
So if anything, arm is more RISC now than it was in 2001.
The only thing I could find is https://www.genymotion.com/blog/just-launched-arm-native-and...
V1 = Slightly tweaked ARM Cortex X1 with SVE ( Used on Snapdragon 888 ) on 7nm aiming at ~4W per Core.
N2 = New Cortex with AMRv9, ~40% IPC improvement over N1 or 10% lower than V1, SVE2, 5nm aiming at ~2W per Core. With Similar die size to N1. ( I fully expect Amazon to go 128 Core with their N2 Graviton )
So in case anyone is wondering, no, it is not Apple M1 level. Not anywhere close.
CMN-700 = More Cores and support of Memory partitioning, important for VMs.
It's apples and oranges comparison to Apple M1 chips (server vs. consumer) but does hint at what's possible with the next generation ARM Cortex "X2" cores, that could appear in next year's flagship smartphones and laptops. A 30-40% IPC jump, partly due to moving to 5nm fabrication process, is huge.
Given the right implementation, namely squeezing more big cores than the current 1-3-4 configuration, it could close the gap considerably with Apple.