Ampere Altra 80-core ARM CPU
servethehome.com
servethehome.com
If I read this right, they reduce their competitors' benchmarks because they have better compilers? Can anyone justify this?
Generally, that is why we prefer to publish "compiler optimized" as best-case performance as well as "GCC" as more of the least common denominator. Both sets of data points are important.
Official SPECint published numbers will not use GCC because the organizations that submit them always want to see the best performance. Ampere used a scaling factor off of published numbers.
If you want to see the impact, we have some numbers from my ThunderX2 review: https://www.servethehome.com/cavium-thunderx2-review-benchma...
You can see the impact clearly there even though that was from a few years ago. Cray has a better performing compiler for ThunderX2 but we did not get to use it due to licensing restrictions.
I hope that helps. The bigger need is for more data since this is one view of performance. There are other needs as well such as FP performance.
When could we expect a review on Altra?
On an Altra review. Great question. I have been bringing it up for some time and live 15 minutes from their headquarters. The invitation is open on our end.
Oh, and don't miss the 3.0 vs. 3.3 GHz.
And even if you do it's irrelevant. Most common large/important frameworks won't compile with proprietary compilers.
Doing Bench's with ICC, XLC or other is hypocritical and often does not reflect anything useful.
Only the HPC world can afford to recompile everything with proprietary compilers and justify the man power to do so. And even so, they already have passed most compute intensive kernels on GPGPU with cuda a long time ago.
In benchmarking you have two choices: publish the real numbers or not. Which option you choose marks you as honest or not.
You can argue about why the numbers for your product are lower in the discussion section of your report. Not in an asterisk.
People can discard with the various SPEC* benchmarks as unrealistic -- they generally are -- but this is a bridge way too far.
The whole point of the SPEC* benchmarks is that you bring out the best of the best and do everything you can to achieve the pinnacle of performance.
And those SPEC* benchmarks are most interesting, and most applicable, to HPC users, who actually do use hardware-specific compilers. Of course another person mentioned that "everything is on GPUs now" (it isn't): If it's on a GPU, then you don't need a 60 core CPU. If it's on a GPU, then SPEC* benchmarks aren't relevant.
Using Intel's compiler or AOCC to compile binaries for normal server use is pretty rare.
https://www.anandtech.com/show/15575/amperes-altra-80-core-n...
Edit: I should note that if you used the X-gene 1 it was very slow, albeit a reliable workhorse for early 64-bit ARM Linux development. These newer chips have far better performance.
I sort of record Applied Micro were doing POWER as well, is that still the case with Ampere?
Marvell bought Cavium which has the ThunderX line. (ThunderX2 being a rather HPC-oriented chip, I think there's a supercomputer already built with it.) Marvell also makes networking-gear-oriented smaller chips (e.g. Armada 8k), one of which is in my little ARM Desktop (MACCHIATObin) :)
NXP (Layerscape) and Mellanox (BlueField) also make network-oriented chips that have around 24 Cortex-A72 cores. NXP's is in SolidRun's newer workstation product.
Meanwhile Amazon bought Annapurna Labs and they make the Graviton (2) for the AWS cloud. This isn't something you can touch physically but it's going to have the biggest impact of all things. This is the real confirmation that Arm servers are legit and the x86/amd64 monopoly is over.
There's also Huawei HiSilicon's Taishan/Kunpeng stuff, which you apparently can buy if you're a serious business, but now it's available in the public Huawei Cloud, but only for the Chinese region it seems??
Oh and Fujitsu is making some epic chip with HBM2 memory and the new Scalable Vector Extensions. But that's only available if you're making supercomputers.
And Nuvia is going to be a thing eventually.. they have not announced anything yet, we have no idea which ISA they are even going to use (could be RISC-V or POWER or SPARC for all we know) but a prominent UEFI/ACPI-on-Arm person is now their VP of Software and is still referring to the Arm ecosystem as "we" https://twitter.com/jonmasters/status/1234734345350369281 :)
And yeah.. press F to pay respects for Qualcomm Centriq and AMD Seattle.
According to techcrunch they have confirmed it will be built on top ARM.
Edit: That is assuming they sort out their lawsuit with Apple.
[1] https://techcrunch.com/2019/11/15/three-of-apple-and-googles...
It had "only" 32 cores. I still find it a lot.
I guess it would be easy to port OpenBSD to the Altra since it boots from UEFI.
On FreeBSD for the eMAG, we've had to:
- ignore a wrong value for UART access width https://svnweb.freebsd.org/base?view=revision&revision=34622... (IIRC Ampere did fix the value in the newer FW revisions)
- restore another register after calling EFI runtime services https://svnweb.freebsd.org/base?view=revision&revision=34699...
- fix some PCIe things we were doing wrong https://svnweb.freebsd.org/base?view=revision&revision=34792... https://svnweb.freebsd.org/base?view=revision&revision=34793...
- fix some memory map things https://svnweb.freebsd.org/base?view=revision&revision=34958...
Not saying that the future has to be multi-die, but if it is not, then it has to be way faster than the cheaper-to-manufacture competition.
If they put 100 cores on every die but only activate 80 of them then that means they can tolerate absolutely HORRIBLE per-processor yields and still make chips that work. Their yields could actually be BETTER than with chiplets because they can afford so many problems.
Not saying that this is true, BTW, just that it's theoretically and practically possible.
This die cost metrics is way overblown and its narrative is too narrowly focused. Especially on ARM Server where unit cost dynamics with ARM IP along with much higher margin on server CPU lower the multi die BOM benefits. And the same definitely does not apply to Graviton, which Amazon owns the whole stack.
Else yield obviously counts, that's what stands in the way of this CPU having more cache or 160 cores, for what it's worth, so it has to count for something obviously. The multiple tiers in every cpu manufacturer line up is also a consequence of yield, so it's very much not a minor element of the equation
More chips per wafer results in more yield, less chips with potential flaws.
> Dr. Hansch’s research at Universitat Munchen for example, shows how as die size increases manufacturers realize an accelerating yield loss and thus accelerating manufacturing cost. Using their model, assuming best-case defect densities, AMD’s small chiplet approach achieves 90% yields vs. Intel’s 30%-40% yields from its large, monolithic die approach. [1]
[1] https://www.barrons.com/articles/amd-stock-can-gain-87-fund-...
>We estimate Intel’s total server die cost at $162 per good server chip while AMD costs about $108 per good server package.
EPYC is 8 Die + an IOD, assuming the massive IOD of ~420mm2 only cost $10 ( That is suggesting GF are selling 14nm / 12nm Wafer at ~$1.2K, even under the AMD WSA with GF would likely not be feasible ), that number suggest the cost per Compute die would be ($106 - $10)/8 = $12, which is again off by quite a bit. Compared to equivalent die size cost of Intel, which is not even accurate because intel do not require the 65%+ Gross Profit Margin from TSMC and it is on an older mature node. And its yield are too pessimistic, we do not have any numbers from Intel on defect per mm2, but taking a guess from Nvidia's massive ~800mm2 die isn't too far off.
So again, let's assume both of those numbers are "relatively" correct. Do you think $62 would matter for an Intel® Xeon® Platinum 8280 Processor with Recommended Customer Price of $10K? Or the same die they are selling at the lower end for $3K?
While it would definitely be good to have those $62 as profits, the reality is on the server market its advantage is relatively minimal.
The bulk of the benefits of Chiplet approach is that you can reuse one or two designs and have it deployed across the whole range of market from Server to Desktop. As design cost increases with Pure Play Foundry such as TSMC this is extremely important for Fabless Design company like AMD, it cost them hundreds of million per design variation. Intel would have that problem down the road but it is mostly migrated for now because all of their design and fabs are in house, and would not cost them as much.
And since the original post was specific to ARM and Server, which has a different market dynamics in cost, you have less R&D as compared to x86 market. In ARM you are paying for IP price on the N1 core and some Interconnect, and those cost of Spread among all ARM players. Hence the conclusion of Chiplet is everything isn't as clear cut and why I said it is too narrowly focused.
I suspect they might have reused these HW blocks from the old Applied Micro X-Gene :D
But now on the new product, it's all Arm Neoverse cores, it's gonna be great.
;)
He would be spinning in his grave. Which would generate an AC current.
There is nothing anywhere near the performance of what I can get right now by buying a $110 motherboard from one of the top six Taiwanese motherboard manufacturers, and a $150 Ryzen 3000 series to socket into it.
Linux and *BSD developers are not going to be shelling out $6000 for a 2U rackmount noisy system that's impossible to operate nicely in their home offices.
Jetson AGX Xavier developer kit - 8 cores, 32 GB, for $700
HoneyComb LX2K mini-ITX motherboard - 16 cores, up to 64 GB, $750
96Boards Developerbox (Socionext) - 24 cores, 32+ GB, $1,200
I'd like a little more than the 4 cores and 4 GB you can get with a Jetson Nano or Raspberry Pi 4.