Supermicro throws its weight behind Arm servers
nextplatform.com
nextplatform.com
But I'm not exactly holding my breath. Gigabyte is another not-entirely-unserious OEM that previously did lots and lots of ARM press releases. And if you look at https://www.gigabyte.com/us/Enterprise/ARM-Server you may actually be impressed.
But, trying to order one of these things is an entirely different matter. Getting pricing, let alone delivery dates, is just impossible, and if you look closer at aforementioned web page, you'll see that the 'show SKUs' button on many models is simply missing.
And this has been going on even before the latest COVID-induced supply-chain crisis, by the way. Getting your hands on useful ARM server hardware has always been plain impossible (unless, I guess, your're AWS, Microsoft or Google). So, while I'm hopeful that this announcement will improve things, I'm not exactly optimistic...
Could be a long wait.
Nuvia promised silicon in mid 2022. It's now January 2023 and no updates from Qualcomm on when they're going to start sampling to OEMs.
We don't have performance numbers to show yet (and not sure if we will be able to publish them when we will), but I can say I am pleasantly surprised with how much cross-compiling is easier than it was ~10 years ago. Especially with zig cc. :)
I tend to treat the desire to cross-compile as a rather strong indicator that the architecture is just not viable for general-purpose computing. Either because it's impossible to get actual hardware at all, or the hardware doesn't have decent specs. In both cases, it's going to be difficult find a business justification for building for this architecture (again for general-purpose computing).
Obviously, architecture bringup is an exception, but we are way past that for aarch64.
But this only applies to a very tiny amount of in-house code which hasn't been tested on aarch64 before.
I'd expect 99.9% of C++ code to just cross-compile and run without issues.
Literally every time in my life I've updated from version A of a compiler to version B of the same exact compiler, I had weeks of change sets required for getting the existing code to build under compiler B.
It's a bootstrapping problem. It takes a lot of infra to run the first build host. We can transition to aarch64 for build hosts if we see the need later.
Having cross-compilation makes things much easier from get-go.
Surely there is some huge win I'm missing here, so if you have figures showing the math you've done, I'm definitely interested.
ARM instances from cloud providers are very appealing to most large-scale enterprise cloud customers, because they represent a significant cost savings at only a moderate upfront investment.
I bought a rk3588 arm widget for $120 ($140 with a nice metal case) with 8GB ram and 8 cores. I installed Ubuntu 22.04 LTS and was able to compile Rust and Go from source. So far I'm impressed. It's pretty fast (about half as fast as my 2015 Xeon) at compiling Rust. It seems generally pretty usable aarch64 platform and bodes well for other aarch64 implementations. I believe the rk3588 requires a custom (not mainline) kernel, but I believe the important bits have been upstreamed and should appear soon ... hopefully.
Neat widget, no fan, cheap, and reasonably fast. Was something like 6-7 times faster than an Rpi4 at compiling rust.
Pretty happy with it, was easy to get ubuntu 22.04 going, but with their image. From what I can tell they use the same not upstreamed kernel on all their images. I've heard progress, in particular with OpenWRT. It does make a decent router, or just a nice mini server. 8GB ram means it can do handle a fair bit. Compiling rust was 6-8 times faster than a Pi4. It's been stable so far, it does require a usb-c power supply, which I had lying around.
It does have a microSD slot.
If in the USA you can get it quicker from amazon, but around $200.
If you have an amd64 desktop or aarch64 server, you can compile to amd64 and aarch64, and powerpc64 with almost the same command. The binary without dependencies can be copied to the server with `scp`. This is the approach we use at ClickHouse[1]. Even more, we always do cross-compilation, even if the host and target architectures are identical. It allows making the build hermetic, not dependent on the environment of the host system.
[1] https://github.com/ClickHouse/ClickHouse/blob/master/PreLoad...
I saw a previous HN comment about this being due to memory bandwidth and cache latency, but I can't seem to substantiate that comment.
That being said, I'd love to have this sort of performance out of servers. Arm released https://www.anandtech.com/show/17575/arm-announces-neoverse-... in 2022, but it's not really available yet anywhere.
Some workloads are obviously faster on the big PC, but I haven't been this excited about a CPU since the Pentium 1.
At the office we all joked that the MBPs were preparing to take off.
- costs more
- plugs into the wall
- cannot easily be moved
- sounds like a helicopter!
My M1 Max MBPs absolutely do not "feel as fast" - but definitely feel proportionately faster for the cost.
I use my laptop as a thin client and have compute jobs running on the PC, which is now a Linux server. Best of both worlds I think.
Same arrangement as here. I wish I could transparently overflow to cloud when the local server gets bogged down by my demands.
I have XRDP setup, but primarily use the machine heedlessly.
I don't know if Apple has released any stats, but it makes a certain amount of sense. The technical reason to put memory and the CPU/GPU on a single package is to decrease memory latency (and possibly boost operating frequency). There's probably also the option to have a really wide memory bus.
Of course, a nice side benefit for Apple is that people can no longer buy 3rd part ram at market rates.
That being said...
> My M1 MacBook is significantly "faster" than my previous Intel 16", even though the per-core performance are roughly similar in CPU benchmarks (small advantage to M1).
The thing about benchmarks is, they're sometimes particular to a specific application, and even data sets. So M1 can be largely tied with an Intel chip across a broad range of benchmarks (e.g. Geekbench), but it can also be a lot faster on your specific workloads.
You've got that backwards, I think. Putting memory on package, or even just soldered on the PCB nearby rather than in removable DIMMs primarily helps hit higher frequencies or lower power at similar frequencies (GPUs and smartphones provide a wealth of examples). There's little or no impact on latency, because most DRAM latency occurs within the DRAM dies or inside the CPU's cache hierarchy rather than on the link between the CPU and DRAM.
The latency does seem improved in the last level cache and main memory, at least compared to AMD's Ryzen 5000 series. Sadly looks like Anandtech has lost the people who did the deep dives and I can't find similar latency numbers for cache and ram for the Ryzen 7000 series.
Apple's power efficiency allows for a 512 bit wide memory system in a laptop, 4 times any x86-64 laptop. I do hope that AMD or Intel ups their game and allows for more aggressive memory systems on laptops. Only similar design I know of is the Xbox series X and PS5, which both use AMD chips with improved memory systems to keep from bottlenecking the iGPU.
The 16" MBP, at least mine, thermally throttles at the drop of a hat. Teams video conf, backups, virus scan, even small builds, etc.
The M1 in comparison is quiet, cool, not sure I've heard the fan going. So sure the performance is similar, the perf/watt is not.
Nominally speaking, its inefficient at that point. You pretty much can make 2 cores fit inside of the M1 core, and each x86 core supports two threads (Would you rather have 4x x86 threads, or 1x M1 core??). I'm intrigued that people continue to find benefits from such a large core (even without SMT / Hyperthreading).
--------
But yes, M1 has huge L1 cache, huge reorder buffers, and extremely wide execution. I'd expect it to win clock-for-clock vs any other core in the market.
But I'm not fully convinced that its the best design / tradeoff. Intel's E-cores + P-cores suggests that modern CPU cores may have become overly big.
Even the E-cores draw more power than the M1/M2 cores.
The interesting thing to me is die-area. Because that's what determines how many cores you get per chip.
But that ruins the original argument... Not that any amount of down clocking can let Intel or AMD get close on the perf/power graph. Besides heat is the main bottleneck on data center density and operational cost.
This is the denialist argument I keep seeing where people want to have their cake and eat it (or perpetually pin their hopes on Zen+1 that hasn't actually shipped).
If Intel could match M1/M2 power/perf they'd do it. But they can't. They can win the perf crown by absolutely burning power or they can get massively trounced to kinda get close on power. Zen is better but still has the same fundamental trade off.
Instead of saying M1/M2 are nothing special I'm super excited to see actual competition in the CPU space for the first time in a long time - competition that has proven a lot of conventional wisdom to be bunk.
??? My argument is, look at server power/performance. AMD EPYC takes the crown. 80-core ARM does not.
Maybe M1 will win, but we've at best got like 8 cores right now. As I stated earlier: M1 cores are huge. I'm not convinced they're better yet, but if Apple wants to make a 32-core M1 or M2 and compare it to an AMD EPYC 64x2 computer, that's when I'll start looking. We will see what they can do moving forward, but I don't expect that this M1 core can scale to a manycore size like EPYC (or Xeon, to a lesser extent).
Benchmarks are always crap, but they're the best we've got. Benchmarks, for now, show that Dual EPYC still is your best computer. At least for the server-scale that Supermicro operates at.
-------
ARM themselves are keen on this. ARM has V, N, and E cores moving forward because they're heading their bets. This Supermicro system is probably an N2 or N1 system? (I haven't looked into it much).
No one else in the world is making cores as large as Apple's M1. Its an aberration, abnormally huge. If Apple fans find it useful then cool but there's other workloads out here. I'm not fully convinced that such a large core like M1 is the best design.
The Intel chip might well spend most of it's time at Max TDP and throttling while the M1 might require unusual circumstances (like maxing out the fast cores, slow cores, matrix multiply (which is outside the core), memory controller, and iGPU simultaneously. There's not really any way to tell from the posted benchmark.
Wonder if the i7-1250 runs in any laptops without a fan.
the process (3nm, ...) too, to get a sense of what fits in an area
[1]https://www.anandtech.com/show/16823/intel-accelerated-offen... [2] https://wccftech.com/apple-5nm-m1-chip-for-arm-macs-2x-perfo...
Today, TSMC has taken the lead, and Apple pays the big bucks to be a prioritized customer above all others.
E cores work the problem the other way, if you narrow the core, you can get about half the work done per core, but in about a fourth the space. That's a win for throughput, but not for latency/interactivity. Drastically different per core performance makes OS schedulers work harder though.
You might get some interactivity benefits depending on your workload and core configuration... One P vs 4 Es is an easy choice, but 4-P vs 16-E, the 4-P is probably going to feel more interactive because UI latency is lower, even though throughput is lower.
The opposite.
A lot of these chips are being designed as "dark silicon", with extremely specific hardware that is rarely used.
When the silicon is dark, they form a heatsink that can absorb excess heat from all other parts of the chip and help regulate the temperature.
I expect more-and-more extremely specific circuitry to be added to future chips, so that more and more silicon can be "dark" and serve as a place for heat to go inside of the microseconds of execution.
I an even imagine spacing functional nodes wider on the silicon, where size and latency limitations allow, to give them more heat-dissipating area.
This would suggest placing E cores between (or even inside) P cores, every cluster of them maybe on a small chiplet.
I wonder if the idea could go as far as having efficient and performant execution units in the same processor, with the reorder buffer issuing instructions to different units according to how far down the instruction stream the result is needed.
At first I thought it would be the close integration of memory, but it seems the consistent instruction size makes the reorder buffers a lot simpler to implement than for x86, so you can get a much larger aarch64 reorder buffer for the same area of an x86 (amd64) one.
A ROB / Reorder Buffer is just RAM from my understanding. Since writes are all out of order, the ROB holds the "future values" of things for sections of code that have already executed (due to out-of-order scheduling), that are waiting for "earlier" code to execute.
Ex: If line#105 has been out-of-order executed, and knows that "MemoryLocationX = 100" happens there... the ROB holds it until line#90 through line#014 are done executing.
-----------------
Consistent instruction size would decrease:
* uop cache (AMD / Intel require a uop cache, where "decoded" instructions stay). This is more optional for ARM... and I'd personally expect this to be the "biggest cost" of x86 instructions.
* Decoder (due to the complexity of x86 instructions, the decoder on those chips is probably bigger).
The ROB / Reorder buffer is "after the decoder", probably just interacting with uops directly. I doubt that instruction set size-or-complexity has anything to do with ROB sizes.
[1] https://www.semianalysis.com/p/apple-m2-die-shot-and-archite...
[2] https://www.hwcooling.net/en/zen-4-architecture-chip-paramet...
Zen3 is 4mm^2 on 7nm process... a process that uses 4x the area per transistor than the 3nm process.
TSMC only just started mass production of 3nm like last week, nothing has shipped yet.
Zen3 was 7nm for sure. So AMD Zen3 ~4mm^2 on 7nm is ~364 million transistors, while Apple's ~5mm^2 on 5nm is ~856 million transistors (https://en.wikichip.org/wiki/File:5nm_densities.svg).
With M1/M2 in the ~5mm^2 area or so, I definitely argue that Zen3 cores are 1/2 the size of M1/M2 cores, transistor-for-transistor.
Its a fat core. Maybe it will work thanks to how advanced processes are getting. Maybe this will encourage others (ie: Intel) to experiment with larger cores as well. Its hard for me to say, but I do welcome the benchmarks.
Also the L2 cache and shared logic make up a larger percentage of M1/M2 per-core at >45%; that's only 20% of the per-core area for Zen 3... if you include LLC that doubles the per-core Zen3 area but only adds 30% for M2...
Point is that M1/M2 and Zen 4 show that the per-core area budget within the same process is now similar across Apple and AMD, not an order of magnitude different. It used to be an order of magnitude different, like back on 32nm Apple A6 was about 8 mm^2/core and Sandy Bridge was 18.5 mm^2/core, or 30 mm^2 including LLC that A6 didn't have.
AMD's L2 cache is 1MB on Zen4, and L3 cache is like 4MB. (and L2 cache compares to Apple L1 cache, while L3 AMD cache is LLC / Comparable to Apple's L2 cache).
AMD's L2 cache is per-core. AMD's L3 cache is decentralized last-level, is 32MB for 8-cores (equivalent to Apple's 16MB for 4 cores). Except... AMD's cores have 2 threads on them while Apple only has 1.
I think my overall point is clear: Apple's cores are abnormally large. AMD / Intel have smaller cores (and larger caches). This is _despite_ shoving 2-thread per core on AMD/Intel through SMT or Hyperthreading.
Remember that only one thread gets that HUGE core on Apple. Its very, very unusual. Even POWER10 (which has oversized cores) allows 8x SMT (8 threads per core) to compensate for its oversized nature.
And the M1's P-core was 2.281 mm^2.
Hyperthreading barely costs any area, but anyway I guess you can say that thanks to that plus the clock speed advantage, Zen4 gets like 15% more performance per mm^2 than M2 P-cores? That's not a massive improvement by any measurement.
(as an aside: if annotations of Zen4 I've seen are correct, its branch predictor has almost as much SRAM as the entire µop+L1i+L1d caches. Which... actually I can completely believe of TAGE)
If you cut out L2 cache, the Zen3 / Zen4 core shrinks significantly.
As per the Zen4 article you had:
> The L2 cache in the cores has increased from 512 kB to 1MB, which also increases the occupied area a bit, but the cores are still smaller overall than Zen 3 on 7nm thanks to the 5nm process. The area including L2 cache is 3.84mm²
The 3.84mm^2 figure _INCLUDES_ 1MB of L2 cache. If you wanna cut that out, doing so will damage your own argument, as the Zen4 core will shrink rather dramatically. (Especially with that "unshrinkable SRAM" argument you're trying to make).
-----------
Look, I don't even know where you're going with this. It shouldn't be a surprise to anybody that a 8x wide M2 core with like 800-entry reorder buffer and 600-entry register file will be bigger than a 6x wide AMD Zen4 core with like 400-entry reorder buffer and like 300-entry register file.
M2 was designed to be big, fat, and wide in execution. That's just how it works. And its a very interesting (arguably brilliant) tradeoff. But if you look at the damn chip, its just bigger. That's what happens when you add more stuff to a core, the core gets larger.
AMD on the other hand, is narrower (especially on a per-thread basis: 2-threads fit on this smaller core), and instead spends way more transistors on L2 cache. Maybe _YOU_ don't like the tradeoff (Sure, I agree that AMD's L2 cache is 15 cycles latency), but maybe throughput is more important and you're overly focused on unimportant / hypothetical latency issues (the entire L2 cache can be accessed at full throughput IIRC).
At the end of the day, we gotta get the devices and then benchmark them with real programs to really see what the sum of all these tradeoffs are. But I don't think there's much argument to be had here that the M1/M2 Apple cores are just bigger. I mean... we know the buffer sizes. We all know Apple's buffers are just bigger.
-------
https://pbs.twimg.com/media/FbXiXZeaAAA2reZ?format=jpg&name=...
Look at this. As I stated before, the biggest "penalty" to the AMD Zen4 core is the uop cache (which is unnecessary in the Apple chip). You can just... look at the damn die shot.
If you want to argue about legitimate space-saving ability of ARM systems, focus on _THAT_ part of the chip. You're talking about all sorts of things that aren't actually helping your side of the argument.
> You pretty much can make 2 cores fit inside of the M1 core
My entire point has simply been debunking this. I'm pointing out that Apple, Intel, and AMD have similarly large area budgets for their big cores. Like, looking at actual chips produced on the same TSMC processes, you cannot fit two Zen4 cores inside of the space taken by one M1 or M2 P-core. You cannot fit two Zen2 or Zen3 cores within the space taken by one A12Z P-core. All of them have a somewhat similar per-core area budget, with difference in L2/LLC cache tradeoffs being the biggest differentiator in area.
And yes, even outside of cache they make different tradeoffs with what they spend the area on. Zen4 spends area on 512b registers and 256b ALUs, and clocking past 5GHz. Apple spends it on scalar resources and deep reordering. I'm not arguing that one tradeoff is universally better than the other, just that they end up similarly big.
Since you brought up cache throughput, Zen4 does 32B/cycle between L1 and L2 [1]. Anandtech measured M1's L2 cache throughput at about 440 GB/s across the 4 P-cores [2], which works out to 34B/cycle/core. Which sure sounds like the same per-core throughput to me.
> You can just... look at the damn die shot
I... did? That's how I estimated L2 cache and tags at 28% of Zen4's 3.84mm^2. Do you believe that that is incorrect, or that the M1 and M2 P-core areas of 2.281-2.756 mm^2 quoted by Semianalysis are incorrect?
[1] https://chipsandcheese.com/2022/11/08/amds-zen-4-part-2-memo...
[2] https://www.anandtech.com/show/17024/apple-m1-max-performanc...
SRAM did not shrink much on N5, and it won't shrink at all between N5P and N3. Apples cores are almost entirely constructed out of SRAM.
I'm pretty sure Zen3 has more SRAM (aka: L2 cache) than the M1/M2 (128kb + 192kB is a LOT of L1 cache, but its still less SRAM than what AMD is stuffing into its cores).
Even with all the extra register files + ROB buffer (also SRAM), the M1 just ain't getting close to 512kB L2 alone (plus all the L1 I$ and L1 D$, and uOp cache, and ROB and Register Files and 256-bit AVX registers on the Zen3).
If anything, bringing up the Apple L1 cache vs AMD L1/L2 per-core caches just emphasizes how big Apple's logic units are in comparison.
I’m going to some personal benchmarks for all all 3 machines plus my new Ryzen 7950x. It still amazes me the fanless M1 Air is nearly as fast as my old desktop and laptop.
i watched this long time ago and i have a new appreciation for "my phone", not to speak of other general computing devices like laptops or desktops
Seriously. The thermals the the last couple of intel apple laptops are atrocious. And I mean from a design standpoint point. With shitty coolers and fan arrangements etc. its kinda wild.
I'd certainly hate to have had an Intel Mac that was louder and hotter than it already was.
I think this conspiracy theory could go the other way: they let the thermals go nuts in old intels to make the M1s look better.
You can compare Threadripper (4-channel memory), Threadripper Pro (8-channel memory, same as for EPYC servers) and M1 by memory bandwidth.
Example 1: https://www.phoronix.com/review/amd-threadripper-5965wx/5 - ClickHouse on Threadripper Pro with less number of cores but better memory bandwidth is faster than Threadripper with larger number of cores.
Example 2: https://benchmark.clickhouse.com/hardware/ various servers, while Aarch64 servers have good places, the top results are from EPYC servers. Never EPYC servers have 12-channel DDR-5 memory - sounds like a heaven for ClickHouse :)
It's clear that Apple and Amazon have their own chips, so will the SM hardware just be sold to Amazon? (I assume Apple has no use for them.)
Early on there was a rush to get everything supported (aarch64 compiled packages mostly), but now everything works just fine.
We just swapped some c6i.2xlarge servers for c7g.xlarge (graviton 3, half the size of the intel servers) and utilization for our workload (sidekiq jobs, in this case) bumped up maybe 20-30%. The performance is quite good.
No issues: if your software supports it, do it.
This stuff is less of a problem if you're just using a few dozen of these things at most. But at scale commodity hardware becomes a nightmare.
It's why most hyperscalers find it more cost-effective to just build their own (Amazon, Google, Meta, others)
That's not a Supermicro-specific problem, though. In fact, the Oxide Computers folks are predicting their whole business approach on being able to "control and support" most every piece of firmware that they ship to users. It's not easy, and it's far from standard in the industry.
And by "hyperscalar" I mean "the ODM who actually designed the system".
Both Microsoft and Meta née Facebook hyperscaled their data centers using the AST1050.
All of the various Open Cloudserver specs call for BMCs: https://www.opencompute.org/wiki/Server/SpecsAndDesigns-old
Things haven't changed that much from these specifications.
Hyperscalers are more likely to want to get OCP equipment or something similar; which might end up being something they bid out to supermicro, but it won't look like the retail servers. As spamizbad notes, the BMC/IPMI firmware is an issue too; it's mitigatable, but something that can run software the owner controls is much preferred (OpenBMC looks nice), if you have the scale to demand it.
I reported a security issue about their BMC / IPMI to Supermicro, and Supermicro decided it wasn't a security issue, so they weren't going to do a thing.
There's no place in this world for companies that choose to wait for active exploitation before doing something.
(edit: link-to-highlight only works in Chromium-based browsers. everybody else, use Ctrl+F)
Oracle even has a free forever tier with ARM: https://www.oracle.com/cloud/free/
Pair any two of those and it strains credibility for anyone that deals with Oracle. All three?
A real shame, as they could offer SPARC, if they hadn't decided the future was x86 after acquiring Sun.
At least IBM offers POWER and Z.
And the few vendors probably can't build enough of them because hyperscalers are syphoning all the fab output to their own datacenters.
Arm was launched in 1985. It has had an extraordinary run. Supermicro tends to stick with tried and true so maybe Arm would be the most cost effective for them. However as you said the instruction set is not the deal breaker any longer - but expensive licenses and proprietary IP can be.
Most ARM server vendors quit before they even had a product on the market.
I have them on my mental blacklist since the reveal they were injecting backdoors into their hardware.
I also dont trust shell games under the label of "changing suppliers" when supposedly the hardware backdoors were added on US soil.
Most references happened in May 2019, but theres one from 2021:
- https://www.bloomberg.com/features/2021-supermicro/?leadSour...